The mechanics behind the impact-collapse latch, the family corroborator, and day-one severity — with the real run-to-failure data — and an honest coverage matrix against the full analyst training shelf: the published SKF CM5003 diagnostic guide and ISO 18436 Cat-II/III analyst coursework.
A rolling-element bearing near end of life goes through a counter-intuitive phase. A young
spall has sharp edges — every ball strike is a clean metallic impact, and the
peak-hold impulse channel (pv: the peak of the
>1 kHz high-passed acceleration)
reads high. As the spall widens and its edges round over, the strikes soften:
the impulse channel falls. On a trend chart the machine looks like it is recovering.
What actually rises is the broadband floor — damage is now everywhere on the
raceway, so energy smears across the spectrum instead of concentrating in discrete impacts.
Days later the bearing seizes.
The analyst courses teach this as the Stage-4 signature: impacts collapsing while the floor rises. Both halves are now in the engine, and both had to be measured in before activation:
High-pass the raw window above 1 kHz, take the absolute peak → pv in g.
Also take the broadband RMS of the same window → the floor.
Both sides of every comparison are 5-scan rolling medians. One startup spike cannot arm the latch; one quiet scan cannot fire it. (Without this, FEMTO's oscillating natural degradations fired at 4–24% of life.)
The running max of the median must exceed the field alert (3 g inner-race row at plateau speed, speed-scaled — §3). A machine that never got loud can never "collapse".
Fire when the median falls to ≤ half of its armed running max and the
median floor now is ≥ the floor recorded when that max was set. A run-in bathtub set its
peak on a high floor and quiets down — the gate blocks it. A widening spall
collapses on a rising floor — the gate passes.
Sticky by design (measured flicker: 4–7 release/re-fire cycles per timeline without it). Persisted in NVS on the chip per bearing-id — survives power cycles; a bearing replacement gets a new id and fresh state. Verdict effect: DEVELOPING-capped — the latch alone escalates, never condemns.
Textbook case: impulse and floor both climb, then the impacts collapse while the floor keeps rising — the latch fires at 0.93 of life, before functional failure.
Run-in vibration set the peak early on a HIGH floor; the mid-life quiet looked like collapse and fired at 0.16 of life. The floor at that moment was falling — gate blocks it. The same gate then passes the genuine end-of-life collapse at 0.95, which the plain latch had already burned itself on. Across all 12 replayed timelines (IMS, XJTU, FEMTO): zero false fires, real fires at 0.91–0.96 of life.
Detection says "something is wrong"; diagnosis names the component. The analyst method (SKF CM5003, the Cat-II drills) treats a defect line and its modulation sidebands as one signature, because the physics of each family modulates differently:
Envelope spectrum around a named defect line. The sideband spacing is the fingerprint: an inner-race defect rides in and out of the load zone once per rev → shaft-rate sidebands; a rolling-element defect is carried by the cage → cage-rate sidebands; an outer-race defect never moves → a clean, unmodulated line.
The engine's corroborator answers only the conditional question — given that the envelope lane has named this line, does the sideband signature agree? — and it is built to abstain rather than guess:
Median-floor SNR saturates on dense rig spectra (healthy Paderborn floors put "SNR" in the thousands for junk between rig lines). A genuine named line must be ≥ 0.25× the largest line in the whole envelope spectrum — a threshold measured across five labeled cohorts, not chosen. A harmonic-position guard was tried first and removed: measurement showed healthy junk peaks sit closer to the defect position 78% of the time.
Each candidate must be ≥ 15% of its carrier (true modulation measured 20–48%; coincidences ≤ 10%), must not sit on the integer shaft ladder, and a "sideband" exceeding its carrier means the spectrum is too dense to localize — refuse.
BPFO ≡ z·FTF exactly, for every bearing — cage-rate sidebands around an outer-race line are frequency-identical to the cage ladder and can never count there. Same for geometries where BSF/FTF lands near an integer (the CWRU 6205: 5.92). The corroborator knows what it cannot know.
bpfo-line + clean → outer. bpfi-line + shaft sidebands → inner. bsf-line + cage sidebands → rolling element. Anything else → ambiguous. It cannot relabel the envelope lane's call, and it is never invoked without one.
Baselines take days to learn. The analyst shelf carries field severity charts — pk-pk g alert levels per defect family, scaled by speed — that three independent documents agree on. They are now the engine's absolute severity anchor: the family named by the envelope lane picks the row, the tachless speed estimate picks the point on the curve, and a machine gets severity context on scan one.
Alert thresholds vs speed (fault = 2× alert). Below 900 rpm the alert falls as (rpm/900)0.75 — which is why the rail axlebox (754–928 rpm at line speed) sits exactly at the knee, and why the low-speed power law is our operating point, not an edge case. Validated: 0 false alarms on healthy across MAFAULDA (49/49) and Paderborn (72/72); deliberately under-calls small lab-rig damages — it annotates, it does not condemn.
| Dataset | What it tests | Result |
|---|---|---|
| FEMTO (natural run-to-failure) | drop latch + floor gate | fires 0.93–0.96 of life · 0 false fires · bathtub false fire eliminated |
| IMS / XJTU (run-to-failure) | latch arming discipline | never fires unarmed · never on monotonic runs |
| CWRU (6205, labeled) | corroborator, inner leg | inner 12/12 · ball abstains by identity · 0 wrong |
| MAFAULDA (full rig, labeled) | field tables + corroborator + ω² lane | healthy 49/49 clean · outer 72/120 · 0 in-contract wrong · ω² exponent 1.70→3.20 with load |
| MAFAULDA — ω² activation (880 files) | speed-curve evidence, switched live | 133 fires, every one a true imbalance (15–35 g) · 0 false fires on all 547 healthy + misaligned controls · 6 g and most 10 g deliberately left undetected |
| Paderborn (6203, artificial + real) | third-dataset stress test | outer 129/144, 0 wrong · healthy fields 72/72 clean · exposed + fixed the floor-saturation hole |
| Baseline binning (fleet replay) | speed-varying false alarms | healthy FP 4.2% → 1.8% · 103/120 fault detections preserved |
| JNU (seventh dataset — blind) | unseen rig, labels hidden until after the run | 9/9 faults detected from a cold start · 0/3 healthy records misread once the calibration fix landed · fault naming 2/9 — the dataset never states its bearing |
Three of this month's four fixes came from a validation run finding something wrong — the corroborator's prominence gate exists because Paderborn broke the old one. That is the loop working as designed.
Below, every fault family those courses teach, against what the engine actually does — with the reason for every gap, because the reason decides whether it is closable.
| Fault (per the corpus) | Their signature | Status | What we run |
|---|---|---|---|
| Bearing defects CM5003 p.20; 4-stage model; impulse/envelope technique notes |
non-integer multiples 4×–15×, enveloping/impulse channels, 4-stage progression | Deepest | 6-path envelope + kurtogram/SK + MED + cepstrum + squared-envelope, the corroborator, absolute envelope + peak-impulse lanes with sourced severity charts, D2 trending → RUL. Validated on all six datasets. |
| · Stage 4 (end of life) | impacts collapse, floor rises, discrete orders dissolve | Covered (new) | drop latch + rising-floor gate, active on all 6 targets. Remaining half: a standalone broadband-floor detector (planned). |
| Imbalance CM5003 p.16 |
dominant sinusoidal 1×, few harmonics, radial, 90° H↔V phase, amplitude grows with speed² | Covered | the 1×-dominance rule + the cross-axis quadrature check ( their 90° phase check from one triaxial node, no second sensor) + new ω² speed-curve lane: their "grows with speed" made quantitative — exponent 1.70→3.20 tracking 6g→35g load on MAFAULDA. ω² currently advisory pending a combined-evidence activation study. |
| Misalignment CM5003 p.14; over half of all machinery problems |
radial 2× (30–200% of 1×) = parallel; axial 1× = angular; 180° phase across coupling | Family covered | the engine runs their exact 2×/1× severity ladder; angular-vs-parallel is screening-level (axial channel exists on the triaxial node). The across-coupling 180° phase check needs a second sensor — single-node physics, not a software gap. |
| Looseness CM5003 p.18 |
string of integer or ½-integer harmonics 2×–10×, >20% of 1× | Covered, refining | a harmonic-forest rule + sub-synchronous energy. Fractional (½×) harmonic ladder widening is planned — the order features already exist. |
| Bent shaft CM5003 p.19 |
misalignment-like spectrum; distinguished ONLY by axial phase across the machine | Gap — by physics | indistinguishable from misalignment without two-point axial phase. We call the misalignment family and say so; the member needs a second node (rail bogie topology will have them). |
| Cocked bearing CM5003 p.19 |
axial vibration, phase varying around four points on one housing | Gap — by physics | requires multi-point phase on a single housing — a commissioning-survey measurement, not an online single-sensor one. No vendor's single fixed sensor covers this either. |
| Gear defects spectrum table: GMF 20×–200×, sidebands at shaft rates |
gear-mesh frequency ± shaft sidebands, tooth count × rpm | Suspect-level | a gear-mesh suspect tag from harmonic structure; a true GMF lane needs tooth-count metadata per machine (planned — needs per-machine tooth-count metadata, not new DSP). |
| Electrical — rotor bar / eccentricity / stator CM5003 spectrum table |
2× line frequency in vibration; slip-pole sidebands in current | Covered via MCSA | the MCSA lane reads the motor current itself — Thomson Nav severity for broken bars, eccentricity, stator EPVA; 3-gated, declines loudly off-contract. Deeper than the guide's vibration-only check. The vibration-side 2×-line corroborator is planned. |
| Pump vane-pass / cavitation Cat-II coursework |
vane-count × rpm; broadband cavitation noise | Deliberate gap | no validation data in hand → no invented thresholds. Standing policy: a detection lane ships only when a labeled dataset exists to validate it. |
| Resonance / natural frequencies Cat-II/III coursework |
amplification when forcing crosses a natural frequency; bump tests | Defended, not named | speed-binned baselines mean resonance crossings don't false-alarm (that was the 4.2→1.8% fix), and the ω² lane's exponent >2 tail is resonance signature — but no dedicated natural-frequency identification yet. |
| Belts / sleeve-bearing oil whirl Cat-II/III coursework |
belt-length rates; 0.38–0.48× sub-synchronous whirl | Out of scope | belt rates need belt-length metadata; oil whirl is journal-bearing territory — our market (industrial motors, rail axleboxes) is rolling-element. Sub-synchronous energy is already a feature, so whirl would at least raise a flag. |
| Technique | Status | Ours |
|---|---|---|
| Overall vibration / ISO zones | Covered | ISO 10816 Class III with kurt-corroborated hysteresis at the zone boundary |
| Time waveform (crest, kurtosis) | Covered | full time-feature set feeds physics rules + the labeler |
| FFT spectrum + harmonic analysis | Covered | order features 1×–10×, sub-sync, sidebands — on the chip |
| Acceleration enveloping | Covered | relative envelope paths + an absolute envelope lane with speed-scaled thresholds |
| Peak-impulse channel (HFD class) | Covered (this month) | pv lane + sourced severity charts + drop latch — the arc this note documents |
| Phase measurement | Single-point | cross-axis phase from one triaxial node (quadrature check). Across-machine phase needs two nodes — planned free with multi-node rail topology. |
| Acoustic emission (ultrasonic) | Rail: resolved · industrial: open | the airborne-ultrasonic mic was dropped (it wasn't ultrasonic); rail spec v3 resolves it with a contact PZT + envelope AFE. Industrial node: the envelope and peak-impulse lanes carry the early-warning load. |
| Multi-parameter monitoring | Covered | the 6-dimension tiered verdict — physics + trending as authority, NN as labeler, OOD as interlock — is multi-parameter monitoring, automated. |
| Speed-curve evidence (ω²) | Live — beyond the guide | the corpus says "amplitude increases with speed"; we measure the relationship across the speeds a machine naturally runs at and test it against the squared law that centrifugal force obeys. Switched from shadow to live after the 880-file study above; runs on all six targets. Not in their handbook as a computed quantity. |
The honest bottom line: everything the corpus can diagnose from one fixed sensor per bearing, we now run or shadow-run — and the three genuine gaps (bent shaft, cocked bearing, across-machine phase) are two-sensor physics, which the rail bogie topology gives us for free. The deliberate gaps (pumps, belts, journal bearings) are policy: no lane without a labeled dataset.
Every number above comes from data we could, in principle, have tuned against. So we ran the engine against a dataset it had never touched — a different laboratory, a different bearing, a different sampling rate — under a harness that reads no labels at all. Twelve recordings: inner, outer, rolling-element and healthy, at three speeds. Every machine started cold, with no history and no learned baseline, which is the hardest condition the engine can be asked to work in. Truth was revealed only after the runs finished.
| What was asked | Result |
|---|---|
| Did it find the faults? | 9 of 9 — every inner, outer and rolling-element record, at every speed |
| Did it cry wolf on the healthy ones? | No. All three read as normal once a deployment fault was fixed — see below |
| Did it name the fault correctly? | 2 of 9. The dataset does not publish which bearing is in the rig |
The third row is the honest one, and it is a physics result rather than a failure. Naming a bearing fault means knowing the bearing's internal geometry — the defect frequencies are computed from it. With the true geometry withheld we ran a catalog assumption, so the engine looked in slightly the wrong places for the defect lines. Detection never depended on that; it comes from impulsiveness, the peak-impulse lane and severity, none of which need geometry. This is precisely why commissioning asks for the bearing designation: it is the difference between "this machine is damaged" and "the outer race is damaged".
On the first pass all three healthy records came back as confident faults. The cause was not the diagnosis at all: a recent model upgrade had been rolled out to the chip and the test harness but missed one deployment path, which was still loading the previous model while being checked against the new model's guard statistics — a mismatched pair that quietly disarmed the very interlock meant to catch this. Fixed, re-run, and the healthy records went clean.
We are publishing this because it is the argument for blind testing. Six datasets of good numbers did not surface it; one unseen rig did, in an afternoon. The guard that failed is the one that flags input unlike anything the model was trained on — and the lesson, now enforced, is that the model and its guard statistics must always ship as a matched pair.