Writing
Do habit trackers actually work?
The honest answer is: partly, for reasons that have little to do with the features most habit apps compete on.
The part with evidence behind it
Self-monitoring — the plain act of recording what you did — is one of the better-supported techniques in the behaviour-change literature. It is not a vague endorsement: it is a catalogued technique with a number. The Behaviour Change Technique Taxonomy compiled by Susan Michie and colleagues, published in Annals of Behavioral Medicine in 2013, is the standard inventory researchers use to describe what an intervention actually does — 93 distinct techniques, sorted into 16 groups — and "self-monitoring of behaviour" is one of them, filed under goals and planning alongside goal setting and action planning.
That the technique has a name and a number matters more than it sounds. Before the taxonomy, two studies could both report that they had used "self-monitoring" and mean different things, which makes comparing them close to meaningless. A shared vocabulary is what allows the question "does this ingredient work" to be asked across a body of research rather than one study at a time.
Asked that way, it holds up reasonably well. Michie and colleagues' 2009 meta-analysis in Health Psychology looked at interventions targeting diet and physical activity and found that those including self-monitoring were more effective than those without it, with the effect strongest when self-monitoring was combined with other self-regulatory techniques such as goal setting and feedback on performance. Harkin and colleagues, writing in Psychological Bulletin in 2016, pooled 138 studies of progress monitoring and reached a compatible conclusion: monitoring progress toward a goal makes attaining it more likely, and the effect is larger when the progress is recorded rather than merely noticed.
The mechanism is not mysterious. Recording a behaviour makes it visible to you, and a lot of the behaviour people want to change is behaviour they are not accurately aware of. Most people cannot tell you how many days last month they actually went for the walk. The estimate and the count differ, usually in a flattering direction.
So: does the act of tracking help? On balance, the evidence says it plausibly does. That is a real finding and it is the one habit apps are implicitly citing when they say tracking works.
The part with much weaker evidence
What that literature does not establish is that streaks, badges, levels, or daily push notifications help. Those are engagement mechanics, and engagement is a metric that serves the app before it serves you.
There is a reasonable theoretical worry about them, too. When a behaviour is tied to an external reward — a number that goes up, a chain that must not break — the reward can become the reason for doing it. When the reward stops, so does the behaviour. This is contested territory in the research and the effect sizes are argued about, so it should be held as a caution rather than a proven mechanism. But it matches something most people have experienced directly: the streak breaks, and the habit that the streak was supposed to be protecting quietly ends with it.
The tell is that the streak was never the habit. Missing a Tuesday does not undo eleven weeks of running. It only undoes the number.
Where the evidence stops being about apps
One caveat runs underneath all of this, and it is worth stating before the argument goes further: almost none of that research was conducted on habit-tracking apps. It was conducted on interventions — supervised programmes, paper diaries, structured trials with a person checking in. An app that implements self-monitoring is borrowing evidence from a setting it does not fully reproduce.
The borrowing is not illegitimate. The technique is the technique, and the mechanism — making an invisible behaviour visible — does not depend on whether the record is kept in a notebook or a phone. But the trial participants had something most app users do not: a reason to keep recording after the novelty wore off, and someone who would notice if they stopped. Adherence in a study is propped up by the study.
Which points at the real failure mode. The technique with the evidence behind it only works while you are doing it, and the thing most likely to stop you doing it is friction. That is the honest argument for a tracker being an app at all: not that it is more scientific than paper, but that it is in your pocket, and paper is in a drawer.
The question tracking usually fails to answer
Here is the gap. You track a habit for three months and you now know, precisely, that you meditated on 71 of 92 days. What you still do not know is whether it did anything.
That is the question you started with. Almost every tracker answers a different one — how consistent were you — and then presents consistency as though it were the outcome. It is not. It is the input, and an input is only worth sustaining if it produces something.
Answering the real question needs two things most apps do not do: recording an outcome as well as a behaviour, and then actually running the comparison between them.
What a useful answer looks like
It looks like a relationship with a number attached, and an admission when there is not one. "Your mood is higher on days following seven or more hours of sleep, r = 0.31 across 46 paired days" is a finding. "You are on a 12-day streak" is a scoreboard.
It also looks like a system willing to tell you nothing was found — because a system that always finds something is not measuring, it is generating. If an app has surfaced an insight every single week you have used it, that is evidence about the app rather than about you.
This is the thing Crescendo is built around: it correlates habits against sleep and mood, and reports a relationship only when it clears 20 paired days, a coefficient of 0.25 or stronger, and a p-value corrected for the number of comparisons run. Most weeks, early on, that means it reports nothing.
What to do when it reports nothing
A system allowed to return nothing will, sooner or later, return nothing. This is the part people find hardest, so it is worth being concrete about what an empty result does and does not mean.
It does not mean the habit is worthless. A correlation that fails to clear a significance threshold is a statement about the evidence, not about the world: with 20 or 30 paired days, only a fairly strong relationship can clear the bar at all, and a real but modest effect will sit there undetected for months. Absence of evidence arrives long before evidence of absence, and personal data sets are small enough that the gap between them is wide.
It also does not mean the measurement was pointless. Knowing that meditation is not visibly moving your mood is genuinely useful if you had assumed it was — it redirects attention to the variables that might be. The most common quiet finding in personal data is that sleep dominates, and that most of what people attribute to their routine is sleep wearing a different hat.
The practical response to a null result is to change one thing and keep going: extend the window, or make the habit more distinct so that the days it happens differ more from the days it does not. A habit performed at 90% consistency has almost no contrast in the data, and a comparison needs something to compare.
So should you use one
Probably, with clear eyes about which part is doing the work. Record the behaviour, because self-monitoring is the ingredient with support behind it. Record at least one outcome you care about, because otherwise you cannot ever answer the question you are asking. Treat the streak as decoration.
And give it enough time to say something. Twenty paired days is roughly the floor at which a personal correlation is worth reading at all, which in practice means about a month before any of this becomes interesting.