2 SEPTEMBER 2026
An HRV number on its own is close to meaningless. The same value can be unremarkable for one person and unusual for another, and unremarkable for you in June and unusual for you in November. A baseline — your own recent history — is what turns the number into information. Here is how to build one that holds up.
A baseline inherits every quirk of how the data was collected. If you wore the watch overnight for four days and only during the workday for three, your "baseline" is partly a description of your wearing habits rather than your physiology. Posture, time of day and activity all move HRV, so mixed sampling conditions bake mixed conditions into the reference.
Fix that first: one wrist, a band that does not slide, and overnight wear if you can manage the charging. The guide to consistent readings covers the details. It is worth doing before you start counting days, because a baseline built on inconsistent sampling has to be rebuilt.
Apple Watch takes multiple HRV samples a day, on its own opportunistic schedule, and the count varies wildly — a night of good wear might produce several, a busy day with the watch off the wrist might produce none. If you pool raw samples, days with more samples silently count for more, which is not what you want.
The fix is to reduce each day to a single daily value first, then build the baseline from those daily values. Every day then gets one vote regardless of how chatty the sensor was. Days with no usable readings simply do not vote — they should not be filled in with a guess or carried forward from yesterday.
This is the choice that matters most, and it is easy to get wrong.
Suppose six ordinary days and one wrecked one — a late night, a fever coming on, a first hard session back. A mean absorbs that seventh day and drags the whole reference point down. For the next week, every day gets scored against a baseline that has been quietly biased by the one day you would least want in it. A median ignores it entirely: it takes the middle value of the set, so a single extreme reading changes the ordering but not the centre.
The same logic protects you in the other direction. One unusually high day — a long lie-in, a rest day after a deload week — will not inflate your reference and make the following week look worse than it was.
Medians are less familiar than averages, so the intuition is worth stating plainly: an average asks what is the total, spread out? A median asks what does a typical day look like? For a noisy signal with occasional wild days, the second question is the one you actually want answered.
The honest answer is that it depends on how noisy your own data is, but the shape of the timeline is consistent:
Note what is not on that list: a day count after which HRV becomes precise. It does not. A longer history makes the reference more stable; today's individual reading stays as noisy as it ever was.
Not medical advice. Calmline is a consumer wellness app. A baseline is a personal reference point, not a clinical measurement — it does not diagnose, screen for, treat or predict any condition. Take health concerns to a clinician.
A baseline fixed at setup slowly stops describing you. Your typical HRV moves with age, fitness, sleep patterns, season and circumstance. A reference from six months ago will eventually tell you that you are permanently stressed, or permanently fine, when in fact your normal has simply relocated.
A rolling window solves this: the baseline is always the median of the most recent stretch of days, so it walks forward with you. Calmline uses a seven-day rolling median for this reason. The trade-off is real and worth naming — a rolling baseline adapts to genuine change, but it also adapts to change you might not want it to absorb. Two weeks of poor sleep will drag your baseline down until "normal" includes them, and then a merely-bad day stops looking unusual.
That is the strongest argument for keeping longer history around and occasionally looking at it. The rolling median answers "is today unusual for me lately?". The longer record answers "has lately itself drifted?". You want both questions available, and they need different windows.
With a baseline in place the arithmetic is simple. Calmline publishes it openly on the home page:
Today below your own median produces a higher score; at or above it, the score sits near zero. There is no model to trust and nothing hidden — the whole calculation is one line.
What a high score means is narrower than people assume. It means today's HRV is suppressed relative to your own recent typical value. It does not say why. HRV has no way to distinguish a difficult week from a head cold, a late glass of wine, an early flight or a hard session two days ago. Anything claiming to separate those from HRV alone is guessing.
Used properly, the baseline turns a number into a question: something is loading your system today — what is it, and does today's plan still make sense? That is a genuinely useful prompt, and it is the whole job. For the reading side of this, see how to read a stress score against your own baseline, and for why nobody else's baseline helps you, why you shouldn't compare your HRV to anyone else's.
Calmline builds this baseline on your device from the HRV Apple Watch has already recorded, using a seven-day rolling median of daily values, and refuses to show a score with fewer than three usable days. Everything is computed on the watch and iPhone — there is no account, no server and no analytics, because the app contains no networking code at all. Details are in the privacy policy, and you can export every reading as CSV or JSON with the formula included, so your history remains yours and remains readable without the app.
If your readings are still sparse or erratic, start with getting a consistent reading — the baseline is only ever as good as what goes into it.
Calmline builds a rolling median from your own readings and scores today against it, entirely on device. It isn't on the App Store yet — the waitlist is the way in.