Skip to article
//The Lab · TP-690 · Coach engineering

A plan built
around you.

A goal, a calendar and a heart-rate trace tell us different things. The new Coach brings them together into varied training weeks — and learns without mistaking one tired evening for lost fitness.

24 September 2026 Evidence: source audit + software verification Status: implemented · private testing

The rebuild started with a disappointing run. The workout felt too cautious, the warm-up looked more complicated than the work, and it was hard to see how the session belonged to a larger goal. A treadmill can follow its instructions perfectly and still deliver an unsatisfying training experience.

The brief was practical: understand the runner, ask what they want to achieve and how much time they have, then give each run a clear purpose. Easy days, long runs, progression, intervals and hills should appear when they fit the plan. Heart-rate control and treadmill control should make that plan easier to follow.

From evidence to the next runImplemented workflow
  1. 01Understand

    History, confirmed routine
    and a measured profile

  2. 02Plan

    Goal, available days
    and time limits

  3. 03Run

    A clear workout,
    translated to the treadmill

  4. 04Review

    Execution, effort
    and comparable evidence

The next review can inform the plan or refine a starting speed. It does not turn every new heart-rate reading into a new training decision.

Start with the runner, not a template

With permission, Coach reads a 90-day overview of running workouts from Health: when they happened, duration, distance, average heart rate, indoor or outdoor status, and source. That provides history. It does not turn an average pulse from a hilly run into a precise estimate of flat-treadmill speed.

For the routine suggestion, it looks at the previous 28 completed days. Local recordings and imported workouts are checked for duplicates, including the app's own exports to Health. The four weekly running-time totals are sorted, and the middle two are averaged. The runner confirms or corrects the result.

One exceptional week should not become your routineIllustrative data · not a runner's history
100 minWeek 1
120 minWeek 2
140 minWeek 3
220 minWeek 4
Suggested weekly routine130 min(120 + 140) ÷ 2
These fictional weeks are already in ascending order. The middle-two average is 130 minutes; a simple four-week average would be 145. The unusually large week has less influence on the starting point.

The assessment supplies a different kind of evidence. Its current protocol uses four-minute speed stages, leaving the first two minutes of each stage out of the steady-state estimate. It checks coverage and whether heart rate has settled, then fits a line through at least three usable speed–pulse points with enough spread.

Think of the line as heart rate = slope × speed + intercept. Reversing it gives an estimated speed for a target heart rate, with uncertainty retained. This is a submaximal field estimate. It does not measure lactate threshold or VO₂max. The runner's resting and maximum heart-rate basis still matters.

A target time is a demand, not a measurement

The goal has a distance, a date and an optional finishing time. The current input range is 1–100 km, including exact distances such as 42.195 km. The planner handles distance continuously: a small edit around 30 km should not suddenly produce a different kind of athlete.

Distance divided by time gives the goal's required average speed. It is useful for shaping the work, but the requested time never makes the measured profile faster. Workout speeds are bounded by that profile.

What does the goal ask of you?Interactive arithmetic · not an ability prediction
Average pace5:00 /km
Average speed12.0 km/h

50 minutes ÷ 10 km = 5:00 per km.

This small calculator shows only the arithmetic of a goal. Coach also needs your measured profile, routine, calendar and available preparation time to build a plan.

Build backwards from the goal, forwards from today

Coach asks for the current weekly routine, a maximum weekly budget, available weekdays, each day's time limit and a preferred long-run day. Available days are choices: offering five does not require five runs. The usable week cannot exceed either the weekly limit or the combined limits of the selected days.

Working backwards, the goal's distance and estimated event duration suggest a long-run requirement and a peak week. Working forwards, the current routine, progression rules and calendar constrain what can actually be scheduled. When the two do not meet, the shortfall becomes a preparation constraint.

For a concrete example, the current marathon rule starts with a 150-minute long-run reference. If the estimated event duration is 4 hours 30 minutes, five-eighths of that is 168 minutes 45 seconds. The requirement is the larger value, capped at four hours: min(240, max(150, 168.75)) = 168.75 minutes. That is a planning demand, not next Sunday's prescription. The build-up and available time still apply.

Weeks move through foundation, build and specific work, with a lighter week in each four-week cycle and easing before the event. Normal-week growth has a 10% ceiling in the current policy. These numbers are explicit product rules, not a claim that one progression formula is optimal for every runner.

Variety has a job to do

Workouts come from structured recipes. The event, training phase and available time influence which recipe fits. Longer-event specific work leans more towards sustained tempo; shorter-event work can use more intervals. Harder sections spend the same weekly budget as everything else.

Easy & recovery

Steady work in a heart-rate band.

Long

Time on feet, sometimes with a faster section.

Tempo & progression

Sustained work or a deliberate rise in pace.

Intervals & fartlek

Faster repetitions with recovery between them.

Hills

Prescribed speed and native incline levels.

The workout presents one logical warm-up and one cool-down block. The executor handles the smaller machine transitions. This keeps the instructions readable without pretending the belt can jump instantly between loads. Warm-up length still depends on the recipe.

Flat easy and long sections can regulate speed against a heart-rate band. Faster repeats and hills use prescribed pace or incline, with pulse supervision still present. In the current hill recipe, the incline request is native level 3, subject to the machine's limits. It is not assumed to mean 3%, and heart rate does not continuously move the slope. Execution sections are currently time-based; the distance shown for a workout is an estimate.

One tired evening is not a new fitness level

Morning and evening runs can feel very different. Coach knows when a run happened; the runner can add how fresh they felt, whether they had recently eaten and how hard the session was. It does not infer a full stomach from the clock.

Comparisons stay within matching contexts, including daypart and workout demands. A run marked tired or after a recent meal still counts as training completed, but is excluded from durable pace learning. Missing answers remain uncertainty. Sensor trouble or a controller-imposed limit must not become evidence that the runner got slower.

This led to four separate feedback loops:

01

Review the run

Compare the prescription with qualified execution data and the runner's feedback. A workout average alone cannot prove that each segment met its target.

02

Adjust the workload

Three comparable, unusually hard runs on separate days, spanning at least six days, need objective underperformance too. That can reduce the next plan week's running time by 5%, once for that week. Six such days can lead to a larger replan proposal.

03

Refine the starting speed

Stable, qualified speed–pulse sections from at least two comparable days can refine the starting speed for heart-rate control. Each adjustment is at most 0.2 km/h; the total stays within 10% of the measured anchor and at most 1 km/h.

04

Reassess the goal

Combine the measured starting range, comparable pace evidence and completed training. Report how the goal stands without silently changing the target.

Reviewing after each run therefore does not mean rewriting every future week after each run. Improvement can support the planned progression. Persistent difficulty can justify an adjustment. A single awkward session does neither by itself. Pain and illness reports are handled separately.

How realistic is the goal now?

The first time range comes from the assessment: estimated speeds at the relevant heart-rate band's ends, converted to times for the chosen distance. Later, reliable steady sections can move that range. To establish a trend, the current model requires two non-overlapping windows, each with at least three running days spanning six days, compared within the same context.

A 2% speed change is the current threshold for a material trend. The central speed shift relative to the assessment is bounded at ±10%; scattered evidence widens the time range. Old evidence is marked stale instead of being silently replaced by a more flattering starting estimate.

Pace is only part of the answer. The assessment also checks recent completed running time against the plan, longer-run preparation, available time and the approaching date. A speed estimate that looks promising cannot fill a missing long run.

The result can be on track, a stretch, reconsider or insufficient evidence, with reasons and an evidence-strength label. It estimates what current evidence supports; it does not invent a probability of success or assume that several remaining weeks guarantee improvement.

The hard parts were also ordinary app problems

Caution could be counted twice. Reducing a starting speed for uncertainty and then lowering the reachable target for that same uncertainty could make the workout unnecessarily timid. The rebuild separated that allowance from the independent execution limits. A useful training prescription still has to reach the intended work.

The calendar needed one answer. Today and Weekly now derive upcoming work from the same saved schedule. If Friday's run was completed early, Sunday may correctly be next. Clear completion states and the actual completion date make that understandable; skipped work remains distinct from completed work in the weekly summary.

Opening a screen is not new evidence. A new run, edited feedback, changed settings or a new local day can warrant work. Reopening unchanged data should reuse a valid result. Corrections and deletions invalidate it, and stale background results cannot overwrite newer inputs.

Saving a response should not wait for every calculation. Feedback is saved first; derived analysis is queued and combined where possible. A meal-context edit does not require processing the same physical samples again. When work does block the interface, the loading view names the actual stage instead of showing a made-up percentage.

What this version does — and what the evidence supports

This is an implemented, deterministic planning system in private testing. The same inputs and policy version produce the same decisions; a language model is not inventing the plan. The current workflow uses running history and explicit feedback, not automatic Health sleep, HRV or VO₂max inputs. Support for a distance entry does not establish suitability for every terrain, ultramarathon or individual athlete.

Regression tests, recorded-data replay and simulator walkthroughs check the software's behaviour. They do not establish improved race outcomes. Performance work separated fast durable saves from slower analysis; cold analysis and some interface stalls still need further measurement. We are not claiming an instant interface or a validated race predictor.

Method. Implementation audit of the personalized Coach and app-design work in TreadPilot revision 21cec66, checked on 24 September 2026. Planning and adaptation claims were reviewed separately against their source and regression cases, including duplicate evidence, morning/evening contexts, missing feedback and arbitrary goal distances. Numerical examples here are illustrative calculations, not personal workout data or training prescriptions. The policy thresholds are engineering choices; this article reports how the product works, not the results of a training trial.

Related reading. For a measured hardware result from the heart-rate controller, see The first live run. That earlier run tests a controller, not the effectiveness of this new training plan.

← Back to The Lab