Three ways field data fooled me about durability (1,350 climbs, 33 amateurs)

Not a product post. One link at the bottom, ignore it if you like.

I build race-day pacing from GPX files. Every pacing model I know of, mine included, prescribes the fourth climb of the day at the same %FTP as the first. That’s physiologically wrong: achievable power declines with accumulated work, and durability is a performance parameter that FTP measured fresh doesn’t capture (Maunder et al. 2021).

The literature is solid in professionals and thin in amateurs, so I ran it on real rides. Three of the four things I thought I’d found didn’t survive. Two of them were my own analysis fooling me, and one of those traps will catch anyone doing within-ride fatigue work, which is the main reason I’m posting.

Dataset

Candidate rides (≥3 h, outdoor, power + elevation) 456
Rides with ≥1 climb ≥5 min with power 364
Climbs observed 1,350
Riders 33
Clean climbs ≥10 min (FTP and weight known) 749

Design

Within-ride comparison: each late climb measured against the first climb ≥10 min of the same ride, which removes rider and day as confounders. 464 pairs. Effort is expressed as %FTP, with effort ratio (mean power against the rider’s own best power at that duration over the preceding 90 days) as a second metric. Both point the same way; %FTP is used in the tables below.

1. Field durability is real, and much smaller than lab durability

Δ kJ/kg vs first climb n Δ %FTP (mean) median
0–5 57 +2.6 0.0
5–10 104 −2.6 −1.9
10–15 103 −1.7 −2.6
15–20 66 −3.5 −3.7
20–30 87 −4.4 −3.5
30–45 40 −4.7 −3.2

Global slope −0.13 %FTP per kJ/kg (r = −0.13). In effort-ratio terms, −0.27 points per kJ/kg (r = −0.21). This is the one result that survived every correction below.

Lab work with maximal efforts reports 6.5–12.5% decline after ~14 kJ/kg. I see 3–4% after 15–30 kJ/kg. I don’t think the literature is wrong, I think we’re measuring different things. Field efforts are voluntary: the rider is pacing, not testing. What this measures is a floor of capacity, not capacity. For prescribing pacing that’s arguably the more useful quantity, because it’s what people actually sustain and still finish on. It is not comparable to a post-fatigue TT and shouldn’t be quoted as if it were.

Note the plateau: the 20-30 and 30-45 bins are nearly identical. The real shape saturates rather than continuing linearly.

2. My model was three times too steep, and being right on average hid it

Comparing my v1 predictions against observed power relative to each rider’s fresh baseline (<8 kJ/kg), n = 515: mean predicted 0.964 against 0.969 observed, and monotonic across bins. Calibrated on average, and wrong exactly where it changes a prescription. Where the model says 0.91, riders hold 0.94. The slope was about 3x reality and only the clamp was saving it.

If you’re building anything similar, this is the failure mode to check for. Aggregate calibration tells you almost nothing about the tail, and the tail is the only part that alters what you’d tell a rider to do.

3. Prior intensity looked like the best predictor. Then the sign flipped.

This was going to be the headline, because it lines up with Spragg 2024 and the 2025 systematic review: intensity of prior work, not quantity, drives the downward shift. The first pass agreed, and the ranking was clean:

Prior-work predictor r
raw kJ/kg −0.13
kJ/kg above FTP −0.17
IF of prior work −0.27

Then two numbers stopped me.

The baseline climb’s own effort correlates −0.51 with Δ, which is by a distance the strongest single relationship in the dataset. And the baseline climb’s effort correlates +0.79 with prior IF, because the baseline climb is inside the prior-work window. Prior IF and baseline effort are close to the same variable. A rider who buries himself on climb 1 therefore shows a high prior IF and a large subsequent drop, for arithmetic reasons.

So I re-ran it two ways: excluding the baseline climb’s window from prior work, and then partialling out the baseline climb’s own %FTP, which is the direct test against regression to the mean.

Prior-work predictor r (baseline climb included) r (baseline window excluded) partial r (controlling baseline %FTP)
raw kJ/kg −0.13 −0.13 −0.16
kJ/kg above FTP −0.17 −0.15 +0.06
IF of prior work −0.27 −0.18 +0.23

Excluding the baseline window isn’t enough on its own: −0.18 is still steeper than raw kJ/kg, and I’d probably have published that. The test that matters is the last column, and there the sign flips. Holding the baseline climb constant, riders who rode harder between climbs went on to climb the late climb harder, not softer.

That’s a good-day effect, not fatigue. On a day when the legs are there, everything is above par, including the bit in between the climbs.

Only raw kJ/kg is stable across all three specifications. Both intensity measures move around or reverse.

I’m not claiming the literature is wrong. Spragg and others used controlled prior work; I have people riding bikes as they please. The honest conclusion is narrower and, I think, more useful: with voluntary field efforts the effect of prior intensity is not identifiable. The point estimate travels from −0.27 to +0.23 depending on a defensible analytic choice, and the rider’s intention sits on both sides of the regression. Anyone measuring prior intensity from ride files is at real risk of measuring how hard the rider felt like going that day.

I’d like to know whether anyone has a design that gets around this without a lab.

4. Level doesn’t separate riders, and neither can I separate individuals

Pooled slope by level: intermediate −0.15, advanced −0.10, both r ≈ −0.1. In professionals the level effect under fatigue is well established (Muriel 2021, Mateo-March 2022, Gallo 2022). In this amateur sample it isn’t there.

Per-rider slopes (23 riders with ≥8 clean climbs ≥10 min): median −0.25, SD 0.27, IQR −0.39 to −0.12, 5 of 23 positive.

I read that spread as evidence durability is individual and worth modelling per rider. Then I compared it against the noise in estimating it. Median standard error on one rider’s slope is 0.17, and 0.16 for riders with 20–30 climbs. The RMS of those standard errors is 0.32, against an observed between-rider SD of 0.27. Subtracting gives a negative variance estimate, i.e. the implied true between-rider SD is zero.

Simulating riders who share an identical true slope and differ only by sampling noise, using each rider’s real number of climbs:

IQR positive slopes
simulated, all riders identical (n=20 each) 0.29 14%
simulated, all riders identical (n=30 each) 0.23 10%
simulated, all identical, each rider’s real n 0.25 13%
observed 0.27 22%

Estimation noise alone reproduces the observed dispersion. One residual worth flagging rather than hiding: the simulation matches the IQR closely but under-predicts positive slopes, 13% against 22%, which is about two riders’ worth of excess. Those are plausibly genuine negative splitters rather than unusually durable riders, but I can’t show that.

So the dataset can’t distinguish “durability varies between riders” from “durability is similar and each rider is measured badly”.

Where that leaves it

One usable result: a small, saturating durability derate, roughly 3–4 %FTP by 20–30 kJ/kg, driven by accumulated work and not by its intensity. That’s what I’ll calibrate against and ship. Everything more ambitious is currently unsupported by field data.

Limitations, stated plainly: voluntary effort means these are floors, not capacities; the baseline climb already carries prior work, so low bins are compressed; the model is linear and the truth saturates; climbs are matched on ≥10 min but not on duration or gradient, and the fourth climb of a route is rarely the same shape as the first; correlations are weak (|r| 0.1–0.3), consistent with field literature (Leo 2022); and regression to the mean via the baseline climb may persist in weaker form anywhere baseline effort correlates with what preceded it.

Two things I’d like from this forum

A design, for 3. How do you measure the effect of prior intensity when the rider chooses both the prior intensity and the effort you’re measuring? Fatigue-conditioned maximal curves (best 20 min after >X kJ, WKO style) measure capacity rather than intention, but need riders who actually go deep late. Is anyone computing these routinely, and is the shape stable across a season? More data won’t fix this one. It’s a design problem.

Data, for 4. Long outdoor rides with power and elevation, several climbs over 10 minutes, and ideally a deep history rather than a handful of rides. Resolving differences of 0.1 %FTP per kJ/kg between riders needs a standard error near 0.05, which at the observed noise level means roughly 300 qualifying climbs per rider. Even a coarse answer needs 70 to 100. That makes per-rider durability modelling from voluntary efforts look like a dead end, and I’d like to be wrong. Enough deep histories would settle it.

What you get back: your own slope, where you sit against these 33, and an honest answer on whether your data is deep enough for the number to mean anything. I’ll post aggregate updates in this thread as the sample grows.

Reply here or DM me.


I build Trainara, a cycling coaching app, which is where these rides come from and where this ends up. If you want to poke at it while you’re here, I have 25 three-month codes for people sending data, valid until 12 October. Entirely optional. I’d rather argue about the method.

1 Like

One thing the data can’t tell me, and it’s been bugging me since.

The effect I measured is small: 3-4 %FTP by the fourth climb of a long day. For most of us that’s 8 or 10 watts. Small enough that I’d expect it to be invisible from the saddle. And yet almost everyone I talk to describes the last climb of a big ride as a completely different animal.

So for anyone who got this far: on your last long ride with three or four decent climbs, what actually happened on the final one? Did you hold roughly the same power and it just hurt more, or did the watts drop whether you wanted them to or not?

I’m curious whether the gap between “3% on paper” and “felt like death” is fatigue, fuelling, heat, or simply knowing you’ve already been out for five hours.

Hi @Supercocus ,

very interesting analysis. I have a small N=1 dataset that illustrates several of the problems you describe, especially the distinction between chosen power and available power.

I’m a 57-year-old amateur, ~95 kg, FTP ~300 W. Over this season I have power/HR/cadence data from several 200 km rides, a 300 km brevet and shorter reference rides. I recently went through them specifically from a durability / kJ/kg perspective.

The most useful example is a 215 km brevet:

  • ~8:22 h moving time
  • ~5,440 kJ mechanical work
  • ~57 kJ/kg
  • NP ~205 W
  • largely ridden solo / without much drafting
  • VI only ~1.14

If I simply split the ride into course sections, power looks like a classic durability decline:

Section NP
Start → ~57 km ~223 W
~57 → 132 km ~206 W
~132 → 174 km ~184 W
~174 → finish ~195 W

But I know from the ride what happened: by the first control I consciously stopped riding with others and decided to pace for a comfortable finish.

So the reduction from ~223 W to ~185–195 W was substantially intentional.

The interesting part is that NP actually increased again in the final, hillier section after more than 4,000 kJ of prior work. Cadence also remained basically unchanged. After the brevet, following almost two hours of rest and food, I rode another 36.5 km home and produced roughly 200 W while pedalling for another ~770 kJ.

So a regression of voluntary climbing power against accumulated work would almost certainly attribute more “durability loss” to me than my actual mechanical capability warrants.

That seems to support your point that field efforts provide a floor of capacity rather than capacity itself.

A second signal was much more repeatable: HR at matched submaximal power

Instead of asking “how hard did I choose to climb?”, I also looked at continuous 60 s windows around 190–210 W with similar cadence.

On the same 200 km brevet the HR response was roughly:

  • 0–10 kJ/kg: ~146–149 bpm
  • 10–30 kJ/kg: still mid/high 140s
  • 30–40 kJ/kg: still ~145–147 bpm
  • 40–50 kJ/kg: ~136–140 bpm
  • 50–60 kJ/kg: low/mid 130s

Power and cadence were essentially unchanged.

This was not just one ride.

On a 300 km brevet (~7,270 kJ, ~77 kJ/kg) I saw a similar downward shift in HR at ~200 W, beginning earlier. On another long solo ride I saw it again. On one other 200 km ride I did not see it clearly.

So in my own data there appears to be a reproducible but non-obligatory late-exercise HR attenuation.

Importantly, it is not the same thing as feeling bad.

On the 300 km ride I felt poor for roughly km 80–220. After fast carbohydrates at ~220 km I felt progressively much better and mechanical power came back, but HR at matched ~200 W stayed unusually low.

That made me separate at least three variables:

  1. chosen power,
  2. available mechanical power,
  3. physiological response to a fixed submaximal workload.

Those are not necessarily moving together.

This suggests a possible field design for your prior-intensity problem

For retrospective ride files I agree that I don’t see a clean way around the endogeneity problem. If the rider chooses both prior intensity and the late effort, “prior intensity” partly measures how good they felt that day.

But I wonder whether a semi-standardised field protocol could get closer without becoming a lab test.

For example, within the same rider:

Protocol A

  • accumulate 30 kJ/kg predominantly below threshold / CP
  • then ride a fixed 20–30 min anchor bout at e.g. 65–70% FTP

Protocol B

  • accumulate the same 30 kJ/kg
  • deliberately include a prescribed amount of supra-threshold work
  • then repeat the identical anchor bout

Keep:

  • total kJ/kg matched
  • fueling prescribed
  • cadence range prescribed
  • same trainer or same road/climb if possible
  • environmental conditions recorded

Then compare:

  • HR at fixed power
  • RPE
  • cadence
  • perhaps efficiency factor
  • and, if desired, a short maximal effort after the fixed bout.

The fixed submaximal anchor does not measure capacity, but it avoids the rider choosing the dependent variable. A subsequent maximal effort would then address capacity.

A crossover with the order of A/B randomised would also reduce the “good day” problem.

That is still not a lab experiment, but it seems more identifiable than trying to infer prior-intensity effects from unrestricted ride files.

One more very mundane field-data trap: elevation channels

Three of four long Karoo files I recently analysed had stretches where altitude simply froze for tens of kilometres while power/HR/cadence remained fine.

I only caught it because the route profile was obviously impossible and I could replace elevation with Strava-corrected GPX data.

If climbs are being auto-detected from FIT elevation, this sort of sensor failure could silently alter:

  • climb selection,
  • gradient matching,
  • accumulated climbing,
  • and therefore which efforts enter the durability model.

Probably worth running explicit altitude QC before climb extraction.

My takeaway from the N=1 data

Your distinction between “field durability” and a post-fatigue capacity test feels exactly right to me.

In my case, voluntary late-ride power clearly contains pacing intention.

What seems more stable is the response to a standardised workload after increasing kJ/kg.

So perhaps there are really two useful field constructs:

  1. Behavioural durability
    What power riders voluntarily choose/sustain as work accumulates.

  2. Physiological durability
    How HR/RPE/etc. respond to a fixed external workload after increasing prior work.

And then a third, harder construct:

  1. Capacity durability
    Best/maximal performance after a defined amount and composition of prior work.

Your climb dataset seems very good for (1).

I suspect (2) could be made surprisingly robust with anchor efforts, while (3) probably really does require either prescribed testing or a very deep history of riders who repeatedly go hard late in rides.

Alexander