CNC Machine Tool Alarm Data Acquisition: How It Actually Works
An explainer for engineers and maintenance planners. We cover which signals carry alarm information, how fast they need to be sampled, where data should be filtered, and what the numbers can and cannot tell you.

In this article
- 1
- 2
- 3
- 4
- 5
- 6
Where CNC Machine Tool Alarm Data Acquisition Gets Its Signals
A machine tool does not expose one clean alarm feed. It exposes several. The control publishes alarm numbers and text over a fieldbus or Ethernet protocol. The electrical cabinet carries hardwired signals on relays and contactors. Spindle and axis drives report their own fault codes. A complete acquisition layer reads all three, then timestamps them on one clock.
The control-side feed is the richest. On a typical mill or lathe you can read alarm number, alarm text, active program block, tool number, spindle speed command and feed override. Most controllers expose these through FOCAS, OPC UA, MTConnect or a vendor SDK. The data is already interpreted, so you get an alarm code rather than a raw voltage. That makes mapping easy, but the update rate is often slower than the machine's real behavior.
Hardwired signals matter for events the control does not log. Door interlocks, hydraulic pressure switches, coolant level floats and emergency-stop relay contacts all sit in the cabinet. Wire them into a digital input module and you capture the physical cause behind a control alarm, not just the symptom. This is the layer that tells you whether a spindle overload was mechanical or electrical.
Drive-level data closes the loop. Servo and spindle drives hold current, torque, following error and DC bus voltage. Poll them at 10–50 ms and you can see a torque spike two seconds before the control trips an overload alarm. That lead time is the whole point of monitoring. Without it you only learn about failures after the machine stops.
- 1Control feedAlarm number, text, program block, tool number, overrides
- 2Hardwired I/OInterlocks, pressure switches, coolant floats, E-stop contacts
- 3Drive dataCurrent, torque, following error, DC bus voltage at 10–50 ms
Sampling Rate and Timing: What You Lose When You Poll Too Slowly
Sampling rate decides which faults you can see. A 1 s poll catches a machine that stopped. A 100 ms poll catches a tool break. A 10 ms poll catches the torque spike that broke the tool. Each step down adds cost in hardware, network load and storage. Pick the rate from the failure you need to prevent, not from what the gateway happens to support.
For status-level monitoring, 1 s is enough. You want to know run time, idle time, alarm state and part count. This is the cheapest layer and it answers production questions rather than maintenance ones. Many shops stop here and still get useful OEE numbers.
For alarm forensics, 50–100 ms works well. Axis following error, spindle load and feed override change on that scale when a cut goes wrong. Store the last 30–60 s in a rolling buffer on the edge device so a trigger can dump the window before and after the event. Buffers cost almost nothing and they save arguments about what happened.
For high-speed machining and hard materials, go to 1–10 ms on the drives. Titanium and Inconel cuts can go unstable in a few revolutions. At 10 ms you can correlate a chatter frequency with a specific tooth pass. Below 1 ms you are usually sampling faster than the drive's own control loop reports, so extra rate buys noise, not information.
- 11 sStatus, run time, part count, alarm state
- 250–100 msFollowing error, spindle load, override changes
- 31–10 msDrive torque and current during hard-material cuts
Filtering on the Edge: Turning Raw Bits into Usable Alarm Events
Raw signals are noisy. A single digital input can bounce for 20–50 ms when a relay opens. If you send every transition to the cloud you get a flood of fake alarms. Debounce in the edge device before anything leaves the shop floor. A 30–100 ms debounce window removes most contact chatter without hiding real events.
Analog channels need a deadband. Spindle load drifts with tool wear and material batch, so a fixed threshold trips constantly. Set the threshold as a percentage of the machine's own baseline and require a minimum dwell time. For example: load above 130 percent of baseline for 500 ms or longer. That rule catches a real overload and ignores a single heavy cut.
State machines beat thresholds for multi-signal faults. A tool break usually shows as a torque spike, then a load drop, then a surface finish alarm. Encode that sequence on the edge device and you get a confirmed event instead of three unrelated alerts. Sequence logic also cuts false positives, because a lone spike no longer raises anything.
Keep the raw data. Filtering decides what raises an alert, not what gets stored. A rolling buffer of unfiltered samples lets you re-run a different rule after the fact. Shops that only keep filtered events lose the ability to ask new questions about old failures.
- 1Debounce30–100 ms on digital inputs to kill contact chatter
- 2DeadbandPercentage of baseline plus minimum dwell time
- 3Sequence logicSpike, then load drop, then finish alarm equals tool break
- 4Raw bufferKeep unfiltered samples so rules can be re-run later
Mapping Alarm Codes to Physical Faults
A control alarm code is a symptom label, not a root cause. The same code can come from a worn tool, a slipping belt, a failing drive or a bad cable. Mapping means building a table that links each code to the signals that were present when it fired. Over time that table tells you which code usually means which fault.
Start with the codes that stop production. On most machines that is a short list: servo overload, spindle overload, overtravel, tool change fault, coolant or hydraulic pressure low, and door or safety circuit open. Log the code, the timestamp, and the drive and I/O states in the five seconds before it fired. That window is usually enough to separate mechanical from electrical causes.
Watch the alarm text as well as the number. Vendors reuse numbers across controller generations, so a number that means overtravel on one model can mean something else on another. Store the text string alongside the code. It costs a few bytes and prevents months of wrong conclusions.
Codes also cluster. A spindle overload followed by a hydraulic low alarm is rarely two faults. It is one fault with a second-order effect. Look at time gaps. Events within a few hundred milliseconds of each other usually share a cause; events minutes apart usually do not.
- 1Log the codeNumber plus text, with a timestamp on one clock
- 2Capture the windowDrive and I/O state for five seconds before the trip
- 3Check gapsEvents under a second apart usually share one cause
What This Data Cannot Tell You
Monitoring does not measure tool wear directly. It measures load, current, vibration and time in cut. Those are proxies. You can build a wear model from them, but it needs a calibration run on your material and your tooling. A model trained on aluminum will not transfer to 17-4PH stainless without new data.
It does not measure part quality either. A machine can hold position perfectly and still cut a bad part because of thermal drift, fixture movement or a wrong offset. If quality is the question, add in-process gauging or post-process inspection. Alarm data answers availability questions, not conformance ones.
Retrofitting old machines has hard limits. Controls from the 1990s may only offer a serial port or a set of dry contacts. You can still get useful status and I/O data, but drive-level parameters are often out of reach. Be honest about that before promising a full data platform on a 25-year-old lathe.
Network and security boundaries matter. Machine controls are not designed to be internet-facing. Keep acquisition on a segmented network, and treat the edge device as the only bridge outward. ISO 27001 practice applies here: least privilege, logged access, and no direct inbound path to the control.
- 1Wear is inferredLoad and current are proxies, not direct measurements
- 2Quality is separateUse gauging or inspection, not alarm data
- 3Old controlsSerial and dry contacts only, no drive parameters
Choosing an Acquisition Layer by Machine Age and Failure Mode
Match the layer to the fault you need to catch, not to the newest hardware.
| Machine / goal | Data source | Sampling rate | What you can catch |
|---|---|---|---|
| Modern control, OEE tracking | Control feed over Ethernet | 1 s | Run time, idle, part count, alarm state |
| Modern control, alarm forensics | Control feed plus drive polling | 50–100 ms | Following error, spindle load, tool break |
| High-speed cut, hard material | Drive parameters direct | 1–10 ms | Chatter, torque spikes, unstable cuts |
| Old machine, status only | Dry contacts and serial port | 1 s or event-driven | Door state, cycle start, basic alarms |
| Safety and hydraulic faults | Cabinet digital inputs | 30–100 ms debounced | Interlocks, pressure loss, coolant level |
| Root cause unknown | All layers on one clock | Mixed, buffered | Sequence of events before the trip |
The Practical Trade-off
If you only need production counts, read the control at 1 s and stop. If you need to explain why a tool broke, add drive polling at 50 ms plus a rolling buffer. Do not pay for 1 ms sampling unless you cut titanium or Inconel and need to correlate chatter.
Questions Engineers Ask Next
Do we need to modify the machine control to read alarms?
Usually not. Most controllers already publish alarm numbers and text over an existing Ethernet or fieldbus port. You enable the interface and read it from an edge device.
Hardwired signals are a different story. If you want cabinet-level I/O, you add a digital input module and tap the existing relay contacts. That is a wiring job, not a control software change.
How much data does one machine generate?
It depends almost entirely on sampling rate and channel count. Status-level data at 1 s for 20 tags is a few megabytes per day. Drive polling at 10 ms across six axes produces hundreds of megabytes per day per machine.
The common fix is to store a short rolling buffer on the edge device and forward only events plus summary statistics. Raw data stays local unless someone requests the window around an alarm.
Can we predict tool wear from alarm data?
You can build a proxy model from spindle load, current and time in cut. It needs a calibration run on your material and your tooling before it means anything.
The model will not transfer across materials without new data. Treat it as a shop-specific tool, not an off-the-shelf feature.
What clock should the whole system use?
One clock, and it should be the edge device. Control timestamps, drive polls and digital input events all get stamped on arrival at the edge.
If you stamp on different devices you will spend weeks reconciling a few hundred milliseconds of drift, which is exactly the window that tells you the root cause.
Does monitoring replace scheduled maintenance?
No. It changes what triggers maintenance. Condition-based triggers replace some calendar intervals, but lubrication, filter changes and calibration still follow a schedule.
Use the data to move from fixed intervals to condition-based ones where the failure mode is measurable. Keep the fixed schedule where it is not.
Send Us the Part and the Machine Details
We machine prototypes and production parts on 127 CNC machines, and we quote within 12 hours with a free DFM analysis.
12-hour quote100% inspectionNo minimum order quantity