Lead discovery — diagnose why beginners are churning before anyone designs a fix. My scope ends at insights and recommendations. Redesign and deployment belong to product and design; if I let those creep into my plan, I'd be signing up to a timeline that isn't mine to own.
A direction Product can act on, in three weeks — not a twelve-week academic study.
Before proposing a single interview, I pulled everything already sitting in the workspace: the analytics dashboard, the App Store reviews the CS team had flagged, and an old usability study. The goal of an audit isn't to find the answer — it's to find out what's already known, and exactly where the gaps are so new research spends its budget on the unknowns.
| Metric (beginner flow) | Q2 | Q4 |
|---|---|---|
| 30-day retention | 41% | 28% |
| Avg. sessions, first week | 4.2 | 2.7 |
| Day-1 intro completion | 76% | 61% |
| Set a reminder | 54% | 38% |
| Explore beyond Day 3 | 49% | 31% |
Every step degraded — not just retention. Completion, reminders, and Day-3 exploration all fell together. That pattern says the problem is broad and early, not one broken button.
Useful signal, but self-selected — only certain users leave reviews. A vocal minority, not the silent majority who just deleted the app. I can't size a strategy on 23 reviews.
The only prior study (6 users, moderated, 8 months ago) flagged a confusing session picker — but praised the content. A redesign of that onboarding shipped 6 months ago, squarely inside the decline window. Yet "didn't know where to start" is still the second-loudest review theme. So the redesign may not have fixed the overwhelm — it may have moved it. That's my first thread to pull.
The numbers tell me what and where. They can't tell me why. That gap is the whole reason for new research.
The PM, Rohan, wanted a survey blasted to the cohort this week — fast numbers by Thursday. Tempting. But a survey can only measure answers to questions you already know to ask. We're in pure discovery: I don't have hypotheses worth measuring yet.
So the sequence is deliberate: 5 interviews first to let beginners surprise me and surface the real themes, then a targeted survey to size how widespread those themes are. Interviews explore; surveys validate. Run them in the wrong order and you measure the wrong thing with confidence.
That's mixed methods with a reason — qual to find the why, quant to prove how many.
These aren't conclusions — they're the threads the audit hands me, each tied to a specific metric or theme. The interviews are designed to confirm, kill, or replace them. I'd write them down before fieldwork so I can't quietly bend the findings toward a favourite.
Beginners face too many choices up front and "don't know where to start." The 6-month-ago redesign may have shuffled the confusion rather than solving it.
The default 10-minute sessions feel too long for a true beginner. There's no gentle on-ramp, so first sessions feel like effort, not relief.
Nothing pulls beginners back after Day 1. Fewer set reminders, fewer explore past Day 3, and "forgot about the app" is the single loudest complaint.
The free trial ends before beginners feel enough value to commit — turning a "not yet" into a permanent "no."
One open question the data can't settle: did the redesign cause this (a regression) or did acquisition quality change? I'd flag that for the data analyst to split retention by channel — keeping it out of my qual scope but on the team's radar.
The plan I'd put in front of Jordan and Rohan. The discipline that matters here: research stops at recommendations (design owns the rest), the survey can't start until hypotheses are solid, and there's no "waiting for results" dead time — interview synthesis runs in parallel while the survey is in the field.
Called out explicitly: hypothesis work may finish Monday — so Week 2 kicks off Tuesday, not a hidden slack day.
The fix for "dead time": Wed–Thu aren't spent waiting on the survey — they're spent fully synthesizing the interviews, so Friday's recs start from half a picture, not zero.
Mitigation: grow to ~15 with the growth team (Priya pulls a wider pool by Wed), and over-recruit so a no-show doesn't sink a day.
Contingency: a second short interview round, run alongside a limited survey — with the honest tradeoff that doing both at once is a compromise forced by time, not the ideal.
Mitigation: a hard response deadline, scheduled reminders, and parallel interview synthesis so a lagging survey doesn't stall the sprint.
I caught myself drafting "redesign + deploy" into the plan and pulled it back. Owning a design timeline I don't control is how a researcher loses the room.
Skimmable, recommendation-first, "what do we do Monday." He wants speed and a clear call — and a Slack update every couple of days so he can manage upward.
Method, sample, synthesis, confidence and limitations. She reviews the interview guide before any session runs — rigor is the deliverable she's judging.
If H1 holds, the fix isn't more content — it's removing choice. A single guided "start here" path beats a picker that makes newcomers feel lost.
A 2–3 minute first session lets beginners feel calm before they're asked for 10. Earn the longer sessions; don't open with them.
"Forgot about the app" is the loudest theme. A reason to come back on day two — not just an opt-in reminder — is where retention is won or lost.
If beginners aren't "ready to commit" when the trial ends, the trial is timed to the calendar, not the user. Tie the ask to a value moment instead.
Pair my qual with the analyst's channel-split to confirm whether the 6-month-ago redesign is the regression — or whether acquisition shifted underneath us.
Discovery points the team at a cause and a direction. Whether a change moves D30 is for design to build and an A/B test to confirm — not for research to promise.
Holding the interviews-first line with Rohan mattered as much as the method itself. The move that worked: open with the number he cares about (the 13-point drop), speak from expertise instead of "I feel," and fold his survey questions into my guide so he stayed an ally, not an obstacle.
My first draft had redesign and deployment inside the sprint. That's not research scope. Pulling it back to insights-and-recommendations kept me from being held to a timeline I don't control. A real lesson in role discipline.
When Rohan caught me suggesting parallel surveys after I'd argued against them, the right move wasn't to get defensive — it was to name it as a time-pressured Plan B with a real tradeoff. Honesty under pressure reads as senior, not shaky.
Pre-align with Rohan in a quick 1:1 before the group sync, so the method debate doesn't play out in front of the director. And I'd quota the sample to deliberately include older, less app-fluent beginners — my pool skews young, which is a real blind spot for a beginner-experience question.
I've written the most likely story from the audit, but five real interviews could overturn it. The value of the plan is that it would tell me — and I'd follow the evidence over my own tidy hypothesis every time.
The whole sprint is a case for spending three weeks understanding the problem so the next three months aren't spent shipping a guess.
This case study stops where the real work would begin: a guide built around the four hypotheses, reviewed by Jordan, then five conversations that either confirm the audit's story or rewrite it. The most useful outcome isn't being right — it's handing Product a diagnosis they can trust enough to build on.
So I ran it. Part two plays the full three weeks end-to-end — the interview guide, five beginners, the synthesis, the survey, and the final deck-and-report handoff. One hypothesis didn't survive contact.
Built around the four hypotheses but written not to lead. The spine is a retrospective walkthrough — I take each beginner back through their own first week rather than asking them to theorize. Rohan's drafted survey questions are folded in as probes, so he stayed bought-in. Note the structure: warm-up before anything sensitive, the drop-off moment in the middle where rapport is highest, trial money-talk last.
Guide discipline I held: no leading questions ("was the onboarding overwhelming?" becomes "tell me about a moment you felt unsure"), past behaviour over future intention, and silence after each answer — the second sentence is usually the honest one. Jordan's one note: move the trial question last so money-talk doesn't colour the rest. I took it.
Strong H3 (habit hook) signal — and a sharper version of it: the reminder existed but was generic and mistimed. Not "no reminder," but "wrong reminder." Killed my assumption that just lifting reminder opt-in would help.
Textbook H1 (overwhelm) — and from exactly the user the audit's pool was missing. The redesign added a richer picker; for a low-fluency beginner that's more choice, not less. He never got far enough to have an opinion on session length.
Cleanest H2 (session length) evidence — and it links to H3. The 10-minute default isn't just "too long," it's a commitment barrier that kills the casual re-open. Length and habit are the same wound.
Complicated H4 (trial). The trial clock wasn't the villain — insufficient habit was. He didn't lack value; he lacked proof to himself that he'd stick. That reframes "trial too short" into "no felt streak before the ask."
A new thread the audit never named: intent mismatch. The flow funnels everyone into a generic "beginner journey" regardless of why they came. Reinforces H1 but adds a cause — overwhelm isn't just volume, it's irrelevance.
I clustered every quote into themes, then asked the uncomfortable question of each: how many of the five, and is it the same thing or am I lumping? Two hypotheses got stronger and sharper, one got absorbed into another, and one collapsed.
4 of 5 named "nothing pulled me back." But it sharpened: the problem isn't missing reminders, it's generic, mistimed ones, plus no felt momentum. Maya, Priya and Tom are all really this.
Daniel and Sarah, the two low-fluency beginners. But the cause isn't "too many sessions" — it's choice without intent. The app asks you to self-classify before it's earned the right to. The redesign moved the overwhelm; it didn't remove it.
Real (Priya, glancingly Sarah) — but it's not a standalone cause. The 10-minute default is a commitment barrier that feeds the habit failure. I'd treat it as a lever under H3, not its own workstream.
The one that broke. Tom — my clearest "trial" user — told me plainly it wasn't the clock or the money. It was that he hadn't built enough habit to believe he'd continue. "Trial too short" was a symptom I'd mistaken for a cause. Lengthening the trial would treat the wrong thing. Killing this is the most useful thing the interviews did.
Beginners don't churn because the trial is short or the content is thin — they churn because nothing builds momentum in the first three days, and the app makes them choose who they are before it's earned a single minute of trust.
With real themes in hand, the survey finally had something worth asking. Fielded to 16 recently-churned beginners (Priya's growth pull got us past the thin pool). Small-n, so I read it as directional, not significant — the job is to check the interview themes aren't a five-person fluke, and to rank them.
The trial reason came dead last — consistent with Tom. The interviews didn't just describe the problem, they predicted the ranking.
"On which day did you stop opening the app?"
81% are gone by day 4. The retention battle is won or lost in the first three sessions — which is exactly where to aim the fix.
Qual and quant point the same way. When five conversations and sixteen surveys agree on the ranking and the timing, I can hand Product a direction with real confidence.
Four recommendations, each tied to a finding and a confidence level — and each a direction for design to explore, not a spec I'm handing down. Note what's not here: "extend the trial." The evidence retired it, so I retired it, even though it was the PM's instinct.
The first three days should be a designed sequence with a felt sense of progress — short sessions, a visible streak, a smart day-2 nudge timed to when the user actually meditated, not 8am default. Addresses the 81%-by-day-4 cliff and the strongest theme. This is the headline.
One intent question (sleep / stress / focus) that routes to one obvious first session — not a library to self-navigate. Fixes both the overwhelm (Daniel) and the intent-mismatch (Sarah) the redesign introduced.
Lower the commitment barrier so "not right now" doesn't become "never" (Priya). Medium, not high — it's a lever under recommendation 01, and I'd want design to A/B the default rather than assume the number.
The data killed "extend the trial." If 01 works and beginners build a streak, the paywall lands on proof, not a calendar — solving Tom's real objection. I flagged this as a deliberate non-recommendation so the team doesn't quietly re-add it.
5 + 16 finds and ranks themes; it does not give statistical significance. I reported percentages as signal, not proof.
I talked to people who left. To know what makes beginners stay, the next round needs retained users too.
My qual is consistent with the redesign causing it — but only the analyst's channel-split can rule out an acquisition shift.
"Extend the trial" was the obvious, PM-friendly answer — and it was wrong. Tom's interview retired it. The discipline that mattered was reporting that against my own and Rohan's prior, and making the non-recommendation as loud as the recommendations.
Daniel and Sarah — the two low-fluency beginners I deliberately recruited — carried the entire overwhelm finding. Had I let the pool skew young and app-fluent like it wanted to, I'd have missed the strongest H1 evidence completely.
The raw interviews gave me four themes. The work was realizing H2 and H4 weren't peers of H3 — one folds in, one collapses. Collapsing four findings into one sentence is what made the recommendation rank-able.
Whether the momentum arc moves D30 is for design to build and an A/B test to settle. Research's job was to make sure they build the right thing. On a simulated run, that's exactly where it landed.