V7 engine · technical note · 2026-07-27

How the engine diagnoses — and what the analyst literature taught it

The mechanics behind the impact-collapse latch, the family corroborator, and day-one severity — with the real run-to-failure data — and an honest coverage matrix against the full analyst training shelf: the published SKF CM5003 diagnostic guide and ISO 18436 Cat-II/III analyst coursework.

01 · The machine that "gets better" right before it seizes

Impact collapse: the Stage-4 drop latch

A rolling-element bearing near end of life goes through a counter-intuitive phase. A young spall has sharp edges — every ball strike is a clean metallic impact, and the peak-hold impulse channel (pv: the peak of the >1 kHz high-passed acceleration) reads high. As the spall widens and its edges round over, the strikes soften: the impulse channel falls. On a trend chart the machine looks like it is recovering. What actually rises is the broadband floor — damage is now everywhere on the raceway, so energy smears across the spectrum instead of concentrating in discrete impacts. Days later the bearing seizes.

The analyst courses teach this as the Stage-4 signature: impacts collapsing while the floor rises. Both halves are now in the engine, and both had to be measured in before activation:

Measure the impulse channel every scan

High-pass the raw window above 1 kHz, take the absolute peak → pv in g. Also take the broadband RMS of the same window → the floor.

Median-smooth both, 5 scans deep

Both sides of every comparison are 5-scan rolling medians. One startup spike cannot arm the latch; one quiet scan cannot fire it. (Without this, FEMTO's oscillating natural degradations fired at 4–24% of life.)

Arm only past a sourced alert level

The running max of the median must exceed the field alert (3 g inner-race row at plateau speed, speed-scaled — §3). A machine that never got loud can never "collapse".

Fire on collapse — but only on a rising floor

Fire when the median falls to ≤ half of its armed running max and the median floor now is ≥ the floor recorded when that max was set. A run-in bathtub set its peak on a high floor and quiets down — the gate blocks it. A widening spall collapses on a rising floor — the gate passes.

Stay latched until bearing service

Sticky by design (measured flicker: 4–7 release/re-fire cycles per timeline without it). Persisted in NVS on the chip per bearing-id — survives power cycles; a bearing replacement gets a new id and fresh state. Verdict effect: DEVELOPING-capped — the latch alone escalates, never condemns.

The real data — FEMTO run-to-failure, bearing 2-1

pv peak (g) broadband floor (g RMS) latch fires

Textbook case: impulse and floor both climb, then the impacts collapse while the floor keeps rising — the latch fires at 0.93 of life, before functional failure.

The trap the floor gate closes — bearing 1-2, the "bathtub"

pv peak (g) broadband floor (g RMS) blocked false fire (0.16) real fire (0.95)

Run-in vibration set the peak early on a HIGH floor; the mid-life quiet looked like collapse and fired at 0.16 of life. The floor at that moment was falling — gate blocks it. The same gate then passes the genuine end-of-life collapse at 0.95, which the plain latch had already burned itself on. Across all 12 replayed timelines (IMS, XJTU, FEMTO): zero false fires, real fires at 0.91–0.96 of life.

02 · Naming the fault — what the analyst guides call diagnosis

The family corroborator: line + sideband signature

Detection says "something is wrong"; diagnosis names the component. The analyst method (SKF CM5003, the Cat-II drills) treats a defect line and its modulation sidebands as one signature, because the physics of each family modulates differently:

Envelope spectrum around a named defect line. The sideband spacing is the fingerprint: an inner-race defect rides in and out of the load zone once per rev → shaft-rate sidebands; a rolling-element defect is carried by the cage → cage-rate sidebands; an outer-race defect never moves → a clean, unmodulated line.

The engine's corroborator answers only the conditional question — given that the envelope lane has named this line, does the sideband signature agree? — and it is built to abstain rather than guess:

The line must be prominent, not just above the floor

Median-floor SNR saturates on dense rig spectra (healthy Paderborn floors put "SNR" in the thousands for junk between rig lines). A genuine named line must be ≥ 0.25× the largest line in the whole envelope spectrum — a threshold measured across five labeled cohorts, not chosen. A harmonic-position guard was tried first and removed: measurement showed healthy junk peaks sit closer to the defect position 78% of the time.

Sidebands must be real modulation

Each candidate must be ≥ 15% of its carrier (true modulation measured 20–48%; coincidences ≤ 10%), must not sit on the integer shaft ladder, and a "sideband" exceeding its carrier means the spectrum is too dense to localize — refuse.

Physics identities force abstention

BPFO ≡ z·FTF exactly, for every bearing — cage-rate sidebands around an outer-race line are frequency-identical to the cage ladder and can never count there. Same for geometries where BSF/FTF lands near an integer (the CWRU 6205: 5.92). The corroborator knows what it cannot know.

Confirm or stay ambiguous — never cross-name

bpfo-line + clean → outer. bpfi-line + shaft sidebands → inner. bsf-line + cage sidebands → rolling element. Anything else → ambiguous. It cannot relabel the envelope lane's call, and it is never invoked without one.

0
wrong family calls, every damaged cohort, four datasets (MAFAULDA, CWRU, Paderborn artificial + real)
12/12
CWRU inner-race files confirmed via shaft-rate sidebands
129/144
Paderborn outer-race files confirmed (81/84 artificial, 48/60 real accelerated-life)
abstains
on Paderborn inner (distributed damage has no load-zone modulation) — honest, documented
03 · How bad is it — on the first scan

Day-one severity: the sourced field charts

Baselines take days to learn. The analyst shelf carries field severity charts — pk-pk g alert levels per defect family, scaled by speed — that three independent documents agree on. They are now the engine's absolute severity anchor: the family named by the envelope lane picks the row, the tachless speed estimate picks the point on the curve, and a machine gets severity context on scan one.

Alert thresholds vs speed (fault = 2× alert). Below 900 rpm the alert falls as (rpm/900)0.75 — which is why the rail axlebox (754–928 rpm at line speed) sits exactly at the knee, and why the low-speed power law is our operating point, not an edge case. Validated: 0 false alarms on healthy across MAFAULDA (49/49) and Paderborn (72/72); deliberately under-calls small lab-rig damages — it annotates, it does not condemn.

04 · The evidence

Six datasets, one scoreboard

DatasetWhat it testsResult
FEMTO (natural run-to-failure)drop latch + floor gate fires 0.93–0.96 of life · 0 false fires · bathtub false fire eliminated
IMS / XJTU (run-to-failure)latch arming discipline never fires unarmed · never on monotonic runs
CWRU (6205, labeled)corroborator, inner leg inner 12/12 · ball abstains by identity · 0 wrong
MAFAULDA (full rig, labeled)field tables + corroborator + ω² lane healthy 49/49 clean · outer 72/120 · 0 in-contract wrong · ω² exponent 1.70→3.20 with load
MAFAULDA — ω² activation (880 files)speed-curve evidence, switched live 133 fires, every one a true imbalance (15–35 g) · 0 false fires on all 547 healthy + misaligned controls · 6 g and most 10 g deliberately left undetected
Paderborn (6203, artificial + real)third-dataset stress test outer 129/144, 0 wrong · healthy fields 72/72 clean · exposed + fixed the floor-saturation hole
Baseline binning (fleet replay)speed-varying false alarms healthy FP 4.2% → 1.8% · 103/120 fault detections preserved
JNU (seventh dataset — blind)unseen rig, labels hidden until after the run 9/9 faults detected from a cold start · 0/3 healthy records misread once the calibration fix landed · fault naming 2/9 — the dataset never states its bearing

Three of this month's four fixes came from a validation run finding something wrong — the corroborator's prominence gate exists because Paderborn broke the old one. That is the loop working as designed.

05 · Coverage vs the analyst literature — SKF CM5003 + Cat-II/III coursework

Fault-by-fault: what we cover, what we honestly don't

Below, every fault family those courses teach, against what the engine actually does — with the reason for every gap, because the reason decides whether it is closable.

Fault (per the corpus)Their signatureStatusWhat we run
Bearing defects
CM5003 p.20; 4-stage model; impulse/envelope technique notes
non-integer multiples 4×–15×, enveloping/impulse channels, 4-stage progression Deepest 6-path envelope + kurtogram/SK + MED + cepstrum + squared-envelope, the corroborator, absolute envelope + peak-impulse lanes with sourced severity charts, D2 trending → RUL. Validated on all six datasets.
· Stage 4 (end of life) impacts collapse, floor rises, discrete orders dissolve Covered (new) drop latch + rising-floor gate, active on all 6 targets. Remaining half: a standalone broadband-floor detector (planned).
Imbalance
CM5003 p.16
dominant sinusoidal 1×, few harmonics, radial, 90° H↔V phase, amplitude grows with speed² Covered the 1×-dominance rule + the cross-axis quadrature check ( their 90° phase check from one triaxial node, no second sensor) + new ω² speed-curve lane: their "grows with speed" made quantitative — exponent 1.70→3.20 tracking 6g→35g load on MAFAULDA. ω² currently advisory pending a combined-evidence activation study.
Misalignment
CM5003 p.14; over half of all machinery problems
radial 2× (30–200% of 1×) = parallel; axial 1× = angular; 180° phase across coupling Family covered the engine runs their exact 2×/1× severity ladder; angular-vs-parallel is screening-level (axial channel exists on the triaxial node). The across-coupling 180° phase check needs a second sensor — single-node physics, not a software gap.
Looseness
CM5003 p.18
string of integer or ½-integer harmonics 2×–10×, >20% of 1× Covered, refining a harmonic-forest rule + sub-synchronous energy. Fractional (½×) harmonic ladder widening is planned — the order features already exist.
Bent shaft
CM5003 p.19
misalignment-like spectrum; distinguished ONLY by axial phase across the machine Gap — by physics indistinguishable from misalignment without two-point axial phase. We call the misalignment family and say so; the member needs a second node (rail bogie topology will have them).
Cocked bearing
CM5003 p.19
axial vibration, phase varying around four points on one housing Gap — by physics requires multi-point phase on a single housing — a commissioning-survey measurement, not an online single-sensor one. No vendor's single fixed sensor covers this either.
Gear defects
spectrum table: GMF 20×–200×, sidebands at shaft rates
gear-mesh frequency ± shaft sidebands, tooth count × rpm Suspect-level a gear-mesh suspect tag from harmonic structure; a true GMF lane needs tooth-count metadata per machine (planned — needs per-machine tooth-count metadata, not new DSP).
Electrical — rotor bar / eccentricity / stator
CM5003 spectrum table
2× line frequency in vibration; slip-pole sidebands in current Covered via MCSA the MCSA lane reads the motor current itself — Thomson Nav severity for broken bars, eccentricity, stator EPVA; 3-gated, declines loudly off-contract. Deeper than the guide's vibration-only check. The vibration-side 2×-line corroborator is planned.
Pump vane-pass / cavitation
Cat-II coursework
vane-count × rpm; broadband cavitation noise Deliberate gap no validation data in hand → no invented thresholds. Standing policy: a detection lane ships only when a labeled dataset exists to validate it.
Resonance / natural frequencies
Cat-II/III coursework
amplification when forcing crosses a natural frequency; bump tests Defended, not named speed-binned baselines mean resonance crossings don't false-alarm (that was the 4.2→1.8% fix), and the ω² lane's exponent >2 tail is resonance signature — but no dedicated natural-frequency identification yet.
Belts / sleeve-bearing oil whirl
Cat-II/III coursework
belt-length rates; 0.38–0.48× sub-synchronous whirl Out of scope belt rates need belt-length metadata; oil whirl is journal-bearing territory — our market (industrial motors, rail axleboxes) is rolling-element. Sub-synchronous energy is already a feature, so whirl would at least raise a flag.

Technique-by-technique (Part 1 of the SKF guide)

TechniqueStatusOurs
Overall vibration / ISO zonesCovered ISO 10816 Class III with kurt-corroborated hysteresis at the zone boundary
Time waveform (crest, kurtosis)Covered full time-feature set feeds physics rules + the labeler
FFT spectrum + harmonic analysisCovered order features 1×–10×, sub-sync, sidebands — on the chip
Acceleration envelopingCovered relative envelope paths + an absolute envelope lane with speed-scaled thresholds
Peak-impulse channel (HFD class)Covered (this month) pv lane + sourced severity charts + drop latch — the arc this note documents
Phase measurementSingle-point cross-axis phase from one triaxial node (quadrature check). Across-machine phase needs two nodes — planned free with multi-node rail topology.
Acoustic emission (ultrasonic)Rail: resolved · industrial: open the airborne-ultrasonic mic was dropped (it wasn't ultrasonic); rail spec v3 resolves it with a contact PZT + envelope AFE. Industrial node: the envelope and peak-impulse lanes carry the early-warning load.
Multi-parameter monitoringCovered the 6-dimension tiered verdict — physics + trending as authority, NN as labeler, OOD as interlock — is multi-parameter monitoring, automated.
Speed-curve evidence (ω²)Live — beyond the guide the corpus says "amplitude increases with speed"; we measure the relationship across the speeds a machine naturally runs at and test it against the squared law that centrifugal force obeys. Switched from shadow to live after the 880-file study above; runs on all six targets. Not in their handbook as a computed quantity.

The honest bottom line: everything the corpus can diagnose from one fixed sensor per bearing, we now run or shadow-run — and the three genuine gaps (bent shaft, cocked bearing, across-machine phase) are two-sensor physics, which the rail bogie topology gives us for free. The deliberate gaps (pumps, belts, journal bearings) are policy: no lane without a labeled dataset.

06 · The test we could not tune for

A rig we had never seen, with the answers hidden

Every number above comes from data we could, in principle, have tuned against. So we ran the engine against a dataset it had never touched — a different laboratory, a different bearing, a different sampling rate — under a harness that reads no labels at all. Twelve recordings: inner, outer, rolling-element and healthy, at three speeds. Every machine started cold, with no history and no learned baseline, which is the hardest condition the engine can be asked to work in. Truth was revealed only after the runs finished.

What was askedResult
Did it find the faults? 9 of 9 — every inner, outer and rolling-element record, at every speed
Did it cry wolf on the healthy ones? No. All three read as normal once a deployment fault was fixed — see below
Did it name the fault correctly? 2 of 9. The dataset does not publish which bearing is in the rig

The third row is the honest one, and it is a physics result rather than a failure. Naming a bearing fault means knowing the bearing's internal geometry — the defect frequencies are computed from it. With the true geometry withheld we ran a catalog assumption, so the engine looked in slightly the wrong places for the defect lines. Detection never depended on that; it comes from impulsiveness, the peak-impulse lane and severity, none of which need geometry. This is precisely why commissioning asks for the bearing designation: it is the difference between "this machine is damaged" and "the outer race is damaged".

The blind run found a real bug — in us

On the first pass all three healthy records came back as confident faults. The cause was not the diagnosis at all: a recent model upgrade had been rolled out to the chip and the test harness but missed one deployment path, which was still loading the previous model while being checked against the new model's guard statistics — a mismatched pair that quietly disarmed the very interlock meant to catch this. Fixed, re-run, and the healthy records went clean.

We are publishing this because it is the argument for blind testing. Six datasets of good numbers did not surface it; one unseen rig did, in an afternoon. The guard that failed is the one that flags input unlike anything the model was trained on — and the lesson, now enforced, is that the model and its guard statistics must always ship as a matched pair.