Simulated work scenario · UX Research lead · meditation app

A retention drop,
a skeptical PM,
and three weeks.

30-day retention for the meditation beginner flow slid from 41% to 28% in two quarters. I was pulled in as research lead to own discovery before anyone threw design at it — in a 3-week sprint, with a PM who wanted a survey out by Thursday. This is how I'd audit, plan, and defend that research.

↘ 41% → 28% in 2 quarters 3-week discovery sprint Interviews → targeted survey Stakeholder management Portfolio simulation
✦ how to read this case study
This is a simulated scenario I ran as practice. The real work on display is the discovery thinking — auditing existing data, choosing a method and defending it, and planning three weeks that hold up under stakeholder pressure.
The setup is the work

The brief, the data audit, the interviews-first decision, the sprint plan and the stakeholder push-back all happened in the simulation. That's the core of this study.

Hypotheses = preliminary

They're drawn honestly from the real audited metrics and reviews — but they're starting points the interviews would test, not conclusions.

Findings = projected

The sprint hadn't run yet, so the closing "where this lands" section is where I'd expect the evidence to point — clearly labelled, never claimed as measured.

The Brief

Monday, 9:04 AM. A Slack ping. ✦
JK
Jordan KimRESEARCH DIRECTOR9:04 AM
"30-day retention for new users in the meditation beginner flow went from 41% down to 28% over two quarters and nobody has a clear answer why. I'm pulling you in as research lead — own the discovery phase end to end. Figure out what's actually going on before we throw design solutions at it. I need a research plan, scoped for a 3-week sprint, by Wednesday."
"Heads up — the PM is Rohan. He's already got hypotheses he'll want to push on you. Manage that diplomatically 😅"
✦ what I'm actually being asked to do

Lead discovery — diagnose why beginners are churning before anyone designs a fix. My scope ends at insights and recommendations. Redesign and deployment belong to product and design; if I let those creep into my plan, I'd be signing up to a timeline that isn't mine to own.

A direction Product can act on, in three weeks — not a twelve-week academic study.

THE CONSTRAINTS I'M HOLDING
3-week sprint · plan due Wed
Fast enough for leadership, rigorous enough to trust.
A PM pushing for speed
Rohan wants a survey out now. I have to defend discovery-first.
8 recruitable users (a thin pool)
From the last cohort. Not enough — I'll need to grow it.

First, the Audit

never design new research before reading what exists ✦

Before proposing a single interview, I pulled everything already sitting in the workspace: the analytics dashboard, the App Store reviews the CS team had flagged, and an old usability study. The goal of an audit isn't to find the answer — it's to find out what's already known, and exactly where the gaps are so new research spends its budget on the unknowns.

✦ quantitative · Mixpanel dashboard
Metric (beginner flow)Q2Q4
30-day retention41%28%
Avg. sessions, first week4.22.7
Day-1 intro completion76%61%
Set a reminder54%38%
Explore beyond Day 349%31%

Every step degraded — not just retention. Completion, reminders, and Day-3 exploration all fell together. That pattern says the problem is broad and early, not one broken button.

✦ qualitative · 23 App Store reviews (90 days)
11
"Forgot about the app after the first day"
9
"Overwhelming at first, didn't know where to start"
7
"Sessions too long — 10 min felt like a lot for a beginner"
6
"Free trial ran out before I felt ready to commit"

Useful signal, but self-selected — only certain users leave reviews. A vocal minority, not the silent majority who just deleted the app. I can't size a strategy on 23 reviews.

✦ the timeline clue everyone missed

The only prior study (6 users, moderated, 8 months ago) flagged a confusing session picker — but praised the content. A redesign of that onboarding shipped 6 months ago, squarely inside the decline window. Yet "didn't know where to start" is still the second-loudest review theme. So the redesign may not have fixed the overwhelm — it may have moved it. That's my first thread to pull.

✦ what the data can't tell me
  • Why beginners feel overwhelmed or forget — the story behind the drop-off.
  • Whether the redesign specifically caused it, or acquisition mix shifted.
  • What "ready to commit" means to a beginner — and when that moment is.
  • The emotional + contextual reasons people quietly leave.

The numbers tell me what and where. They can't tell me why. That gap is the whole reason for new research.

The Call

interviews first — and why I held that line ✦

The PM, Rohan, wanted a survey blasted to the cohort this week — fast numbers by Thursday. Tempting. But a survey can only measure answers to questions you already know to ask. We're in pure discovery: I don't have hypotheses worth measuring yet.

So the sequence is deliberate: 5 interviews first to let beginners surprise me and surface the real themes, then a targeted survey to size how widespread those themes are. Interviews explore; surveys validate. Run them in the wrong order and you measure the wrong thing with confidence.

That's mixed methods with a reason — qual to find the why, quant to prove how many.

DEFENDING IT IN THE ROOM
ROHAN PUSHED →
"We have 8 users now. Why not survey today and have data by Thursday?"
I HELD →
"Interviews build the hypotheses the survey needs. Without them we won't know what to ask — we'd risk measuring the wrong thing."
ROHAN AGAIN →
"Aren't the App Store reviews enough to build a survey from?"
I HELD →
"Reviews are self-selected — a vocal minority. Building on them risks optimizing for the loudest, not the typical, beginner."
THE BRIDGE →
I took Rohan's drafted survey questions as inputs to the interview guide — keeping him bought-in without compromising the order.

Working Hypotheses

what the audit suggests — for interviews to test ✦

These aren't conclusions — they're the threads the audit hands me, each tied to a specific metric or theme. The interviews are designed to confirm, kill, or replace them. I'd write them down before fieldwork so I can't quietly bend the findings toward a favourite.

H1Onboarding overwhelm.

Beginners face too many choices up front and "don't know where to start." The 6-month-ago redesign may have shuffled the confusion rather than solving it.

SIGNAL → Day-1 completion 76% → 61% · 9 reviews
H2Session length mismatch.

The default 10-minute sessions feel too long for a true beginner. There's no gentle on-ramp, so first sessions feel like effort, not relief.

SIGNAL → avg sessions wk1 4.2 → 2.7 · 7 reviews
H3The habit hook is failing.

Nothing pulls beginners back after Day 1. Fewer set reminders, fewer explore past Day 3, and "forgot about the app" is the single loudest complaint.

SIGNAL → reminders 54%→38% · explore-D3 49%→31% · 11 reviews
H4The trial expires too early.

The free trial ends before beginners feel enough value to commit — turning a "not yet" into a permanent "no."

SIGNAL → 6 reviews · "ran out before I was ready"

One open question the data can't settle: did the redesign cause this (a regression) or did acquisition quality change? I'd flag that for the data analyst to split retention by channel — keeping it out of my qual scope but on the team's radar.

The 3-Week Sprint

tight, sequenced, no dead days ✦

The plan I'd put in front of Jordan and Rohan. The discipline that matters here: research stops at recommendations (design owns the rest), the survey can't start until hypotheses are solid, and there's no "waiting for results" dead time — interview synthesis runs in parallel while the survey is in the field.

Week 1
Discover — interviews
MON
Draft interview guide (Jordan reviews first) · recruit & schedule
TUE
Interviews 1–2
WED
Interviews 3–4
THU
Interview 5 · debrief notes
FRI
Begin hypothesis synthesis (may bleed into Mon)

Called out explicitly: hypothesis work may finish Monday — so Week 2 kicks off Tuesday, not a hidden slack day.

Week 2
Validate — survey + synthesis in parallel
MON
Finalize hypotheses
TUE
Draft & field targeted survey (n≈15)
WED
Interview deep-dive synthesis (survey in field)
THU
Finish synthesis · monitor responses
FRI
Begin recommendations

The fix for "dead time": Wed–Thu aren't spent waiting on the survey — they're spent fully synthesizing the interviews, so Friday's recs start from half a picture, not zero.

Week 3
Synthesize & hand off
MON–TUE
Finalize recommendations · write up limitations & confidence
WED
Handoff — deck for Rohan (PM), report for Jordan (director)
Scope ends here — design + deployment is product & design's to own.

Risks & Handoff

name the risks before someone else does ✦
✦ risk · thin recruitment pool
8 users isn't enough — and dropouts would hurt.

Mitigation: grow to ~15 with the growth team (Priya pulls a wider pool by Wed), and over-recruit so a no-show doesn't sink a day.

✦ risk · hypotheses unclear after Week 1
Five interviews might not converge.

Contingency: a second short interview round, run alongside a limited survey — with the honest tradeoff that doing both at once is a compromise forced by time, not the ideal.

✦ risk · slow survey responses
Week 2 synthesis depends on responses landing.

Mitigation: a hard response deadline, scheduled reminders, and parallel interview synthesis so a lagging survey doesn't stall the sprint.

✦ guardrail · scope discipline
Research delivers insight, not the redesign.

I caught myself drafting "redesign + deploy" into the plan and pulled it back. Owning a design timeline I don't control is how a researcher loses the room.

✦ two audiences, two deliverables
FOR ROHAN · PM
A decision-oriented deck

Skimmable, recommendation-first, "what do we do Monday." He wants speed and a clear call — and a Slack update every couple of days so he can manage upward.

FOR JORDAN · RESEARCH DIRECTOR
A full research report

Method, sample, synthesis, confidence and limitations. She reviews the interview guide before any session runs — rigor is the deliverable she's judging.

Where I'd expect it to land

the recommendations this sets up ✦
Projected The sprint hadn't run — these are the directions the audit points toward, to be confirmed by the interviews and survey, then handed to design.
DIRECTION 01
Give beginners one obvious first step.

If H1 holds, the fix isn't more content — it's removing choice. A single guided "start here" path beats a picker that makes newcomers feel lost.

DIRECTION 02
Shorten the on-ramp.

A 2–3 minute first session lets beginners feel calm before they're asked for 10. Earn the longer sessions; don't open with them.

DIRECTION 03
Build the Day 2–3 return loop.

"Forgot about the app" is the loudest theme. A reason to come back on day two — not just an opt-in reminder — is where retention is won or lost.

DIRECTION 04
Re-time the trial to value.

If beginners aren't "ready to commit" when the trial ends, the trial is timed to the calendar, not the user. Tie the ask to a value moment instead.

FOR THE TEAM
Settle the redesign question.

Pair my qual with the analyst's channel-split to confirm whether the 6-month-ago redesign is the regression — or whether acquisition shifted underneath us.

HOW I'D PROVE ANY OF IT
Recommend, then test.

Discovery points the team at a cause and a direction. Whether a change moves D30 is for design to build and an A/B test to confirm — not for research to promise.

What I learned

the parts that aren't on a slide ✦
Managing the PM is the job.

Holding the interviews-first line with Rohan mattered as much as the method itself. The move that worked: open with the number he cares about (the 13-point drop), speak from expertise instead of "I feel," and fold his survey questions into my guide so he stayed an ally, not an obstacle.

I scoped myself wrong at first — and fixed it.

My first draft had redesign and deployment inside the sprint. That's not research scope. Pulling it back to insights-and-recommendations kept me from being held to a timeline I don't control. A real lesson in role discipline.

Recover from contradictions honestly.

When Rohan caught me suggesting parallel surveys after I'd argued against them, the right move wasn't to get defensive — it was to name it as a time-pressured Plan B with a real tradeoff. Honesty under pressure reads as senior, not shaky.

What I'd do differently.

Pre-align with Rohan in a quick 1:1 before the group sync, so the method debate doesn't play out in front of the director. And I'd quota the sample to deliberately include older, less app-fluent beginners — my pool skews young, which is a real blind spot for a beginner-experience question.

It's simulated — and the data could still surprise me.

I've written the most likely story from the audit, but five real interviews could overturn it. The value of the plan is that it would tell me — and I'd follow the evidence over my own tidy hypothesis every time.

✦ the meta-point
Discovery isn't slow — it's the thing that stops you from building the wrong fix fast.

The whole sprint is a case for spending three weeks understanding the problem so the next three months aren't spent shipping a guess.

✦ if I ran the sprint for real

The next move is the interview guide — and then letting five beginners tell me I'm wrong.

This case study stops where the real work would begin: a guide built around the four hypotheses, reviewed by Jordan, then five conversations that either confirm the audit's story or rewrite it. The most useful outcome isn't being right — it's handing Product a diagnosis they can trust enough to build on.

So I ran it. Part two plays the full three weeks end-to-end — the interview guide, five beginners, the synthesis, the survey, and the final deck-and-report handoff. One hypothesis didn't survive contact.

Read part two — the sprint, run → ← back to work
D.
simulated case study · made with care · 2026
← Part one was the plan. This is what happened when I ran it.

The sprint,
run end to end.

Three weeks, played out. The interview guide I'd actually field. Five beginners who half-confirmed my hypotheses and half-rewrote them. A survey that sized what mattered. And the handoff — one deck, one report — that told Product what to do Monday. One hypothesis didn't survive contact.

5 interviews · 1 guide Survey · n=16 3 of 4 hypotheses held Simulated run
✦ still a simulation — but now I let it run
The plan said the findings were projected. So I played the three weeks forward — inventing five plausible beginners, fielding the guide against them, and following the evidence even where it broke my tidy story. This is the honest version of "what if I'd actually run it."
Participants are invented

Five fictional beginners, written to be realistic and varied — not cherry-picked to agree with me.

The method is real

The guide, the synthesis moves, the survey design and the way I'd weight confidence are how I'd actually work.

I let a hypothesis die

A run where everything you guessed is confirmed is a run you rigged. One of the four didn't hold.

WEEK 1 · MON

The Interview Guide

45 min, semi-structured, Jordan reviewed it first ✦

Built around the four hypotheses but written not to lead. The spine is a retrospective walkthrough — I take each beginner back through their own first week rather than asking them to theorize. Rohan's drafted survey questions are folded in as probes, so he stayed bought-in. Note the structure: warm-up before anything sensitive, the drop-off moment in the middle where rapport is highest, trial money-talk last.

00 · WARM-UP — 5 min
Lower the stakes
Tell me about the last time you felt stressed or couldn't switch off. What did you do?
What made you download a meditation app in the first place — why then?↳ probe: what was going on in your life that week?
01 · FIRST SESSION — 12 min
Walk me through day one
Open the app like it's your first time. Talk me through what you see, out loud.↳ where do your eyes go first? what would you tap?
When you finished your first session, how did you feel? Was it what you expected?↳ tests H2 — did 10 min feel long?
02 · THE DROP-OFF — 14 min
The day you stopped
Take me to the last day you opened the app. What happened after that?↳ probe: was it a decision, or did it just fade?
In a perfect world, what would have pulled you back on day two or three?↳ tests H3 — habit hook
Did you ever feel unsure what to do or where to start? Tell me about that moment.↳ tests H1 — overwhelm, without naming it
03 · VALUE & TRIAL — 9 min
When it was worth it (or wasn't)
Was there a moment it clicked — where it felt worth your time? When?↳ locating the value moment vs. the trial clock
Did the trial ending affect your decision? Walk me through that.↳ tests H4 — trial timing
Anything I should have asked but didn't?

Guide discipline I held: no leading questions ("was the onboarding overwhelming?" becomes "tell me about a moment you felt unsure"), past behaviour over future intention, and silence after each answer — the second sentence is usually the honest one. Jordan's one note: move the trial question last so money-talk doesn't colour the rest. I took it.

WEEK 1 · TUE–THU

Five Beginners

quota'd for age + app-fluency — tap a face ✦
MA
Maya · 29
DESIGNER · APP-FLUENT
DA
Daniel · 44
TEACHER · LOW-FLUENCY
PR
Priya · 23
STUDENT · APP-FLUENT
TO
Tom · 36
SALES · MID-FLUENCY
SA
Sarah · 53
NURSE · LOW-FLUENCY
Maya — quit on day 4
signed up during a burnout stretch
"Honestly the content was lovely. I just… forgot it existed. There was no reason to open it on Tuesday — nothing was waiting for me."
"The reminder I set went off at 8am, which is the worst possible time. I needed it at 11pm when my brain wouldn't stop."
WHAT IT MOVED

Strong H3 (habit hook) signal — and a sharper version of it: the reminder existed but was generic and mistimed. Not "no reminder," but "wrong reminder." Killed my assumption that just lifting reminder opt-in would help.

Daniel — quit on day 1
never finished the first session
"I opened it and there were all these categories — sleep, focus, anxiety, courses. I didn't know if I was a 'sleep' person or a 'focus' person. I just closed it."
"I'm not great with apps. I wanted it to just… tell me what to do for five minutes."
WHAT IT MOVED

Textbook H1 (overwhelm) — and from exactly the user the audit's pool was missing. The redesign added a richer picker; for a low-fluency beginner that's more choice, not less. He never got far enough to have an opinion on session length.

Priya — quit on day 6
busy, sampled a lot, committed to nothing
"Ten minutes between lectures is a lot. I'd see the session was ten minutes and think 'not right now' — and 'not right now' just became never."
"If there'd been a three-minute one I'd have done it standing at the bus stop."
WHAT IT MOVED

Cleanest H2 (session length) evidence — and it links to H3. The 10-minute default isn't just "too long," it's a commitment barrier that kills the casual re-open. Length and habit are the same wound.

Tom — quit at trial end, day 7
actually liked it — still didn't pay
"I'd done maybe four sessions. I liked it. But paying felt like committing to a whole lifestyle I wasn't sure about yet. So I let it lapse."
"It wasn't the money. It was — I hadn't proven to myself I'd actually keep doing it."
WHAT IT MOVED

Complicated H4 (trial). The trial clock wasn't the villain — insufficient habit was. He didn't lack value; he lacked proof to himself that he'd stick. That reframes "trial too short" into "no felt streak before the ask."

Sarah — quit on day 2
night-shift nurse, wanted help sleeping
"I came for sleep. But the first thing it pushed was a 7-day 'beginner course' — I didn't want a course, I wanted to fall asleep that night."
"It felt like homework. I've got enough homework."
WHAT IT MOVED

A new thread the audit never named: intent mismatch. The flow funnels everyone into a generic "beginner journey" regardless of why they came. Reinforces H1 but adds a cause — overwhelm isn't just volume, it's irrelevance.

WEEK 1 FRI → WEEK 2 MON

What the five agreed on

affinity-mapped, then pressure-tested ✦

I clustered every quote into themes, then asked the uncomfortable question of each: how many of the five, and is it the same thing or am I lumping? Two hypotheses got stronger and sharper, one got absorbed into another, and one collapsed.

H3
The habit hook is failing → CONFIRMED, strongest

4 of 5 named "nothing pulled me back." But it sharpened: the problem isn't missing reminders, it's generic, mistimed ones, plus no felt momentum. Maya, Priya and Tom are all really this.

4 / 5
H1
Onboarding overwhelm → CONFIRMED + re-caused

Daniel and Sarah, the two low-fluency beginners. But the cause isn't "too many sessions" — it's choice without intent. The app asks you to self-classify before it's earned the right to. The redesign moved the overwhelm; it didn't remove it.

2 / 5
H2
Session length mismatch → CONFIRMED, folds into H3

Real (Priya, glancingly Sarah) — but it's not a standalone cause. The 10-minute default is a commitment barrier that feeds the habit failure. I'd treat it as a lever under H3, not its own workstream.

2 / 5
H4
The trial expires too early → DID NOT HOLD

The one that broke. Tom — my clearest "trial" user — told me plainly it wasn't the clock or the money. It was that he hadn't built enough habit to believe he'd continue. "Trial too short" was a symptom I'd mistaken for a cause. Lengthening the trial would treat the wrong thing. Killing this is the most useful thing the interviews did.

0 / 5
✦ the one sentence the synthesis produced

Beginners don't churn because the trial is short or the content is thin — they churn because nothing builds momentum in the first three days, and the app makes them choose who they are before it's earned a single minute of trust.

WEEK 2 · TUE–THU

Now, how many?

interviews found the why — the survey sized it ✦

With real themes in hand, the survey finally had something worth asking. Fielded to 16 recently-churned beginners (Priya's growth pull got us past the thin pool). Small-n, so I read it as directional, not significant — the job is to check the interview themes aren't a five-person fluke, and to rank them.

"WHAT MOST DESCRIBES WHY YOU STOPPED?" · pick all that apply · n=16
"I just forgot / fell out of the habit"
81%
"Didn't know where to start / too many choices"
63%
"Sessions felt too long to fit in"
56%
"What it offered didn't match why I came"
44%
"Trial ended before I was ready to pay"
19%

The trial reason came dead last — consistent with Tom. The interviews didn't just describe the problem, they predicted the ranking.

A QUESTION ONLY THE SURVEY COULD ANSWER

"On which day did you stop opening the app?"

D1–2
D3–4
D5–7
D8+

81% are gone by day 4. The retention battle is won or lost in the first three sessions — which is exactly where to aim the fix.

CONVERGENCE CHECK

Qual and quant point the same way. When five conversations and sixteen surveys agree on the ranking and the timing, I can hand Product a direction with real confidence.

WEEK 3 · MON–TUE

What I told them to do

ranked by evidence, scoped to research's lane ✦

Four recommendations, each tied to a finding and a confidence level — and each a direction for design to explore, not a spec I'm handing down. Note what's not here: "extend the trial." The evidence retired it, so I retired it, even though it was the PM's instinct.

01
Build a 3-day momentum arc, not a content library.
HIGH confidence

The first three days should be a designed sequence with a felt sense of progress — short sessions, a visible streak, a smart day-2 nudge timed to when the user actually meditated, not 8am default. Addresses the 81%-by-day-4 cliff and the strongest theme. This is the headline.

02
Replace "pick a category" with "tell me why you're here."
HIGH confidence

One intent question (sleep / stress / focus) that routes to one obvious first session — not a library to self-navigate. Fixes both the overwhelm (Daniel) and the intent-mismatch (Sarah) the redesign introduced.

03
Make the first session 2–3 minutes, by default.
MEDIUM confidence

Lower the commitment barrier so "not right now" doesn't become "never" (Priya). Medium, not high — it's a lever under recommendation 01, and I'd want design to A/B the default rather than assume the number.

04
Don't touch the trial yet — re-time the ask, not the clock.
EXPLICIT NON-ACTION

The data killed "extend the trial." If 01 works and beginners build a streak, the paywall lands on proof, not a calendar — solving Tom's real objection. I flagged this as a deliberate non-recommendation so the team doesn't quietly re-add it.

WEEK 3 · WED

The Handoff

two audiences, two artifacts, one ending ✦
FOR ROHAN · PM — the deck
Six slides, decision-first
01The drop in one number, the cause in one sentence
02The day-4 cliff chart
03Four recommendations, ranked
04"Why we're NOT extending the trial"
05What to test first & how we'll know
06Verbatims that make it human
FOR JORDAN · DIRECTOR — the report
Full method & confidence
·Audit → interviews → survey, and why that order
·Sample, quotas, and the young-skew I corrected for
·The hypothesis that died, and what killed it
·Confidence per finding (n=5 qual + n=16 quant)
·Limitations: small-n, churned-only, no channel split
·The open question handed to the data analyst
✦ what I said out loud about the limits
Small, directional sample

5 + 16 finds and ranks themes; it does not give statistical significance. I reported percentages as signal, not proof.

Churned-only voices

I talked to people who left. To know what makes beginners stay, the next round needs retained users too.

The regression is still open

My qual is consistent with the redesign causing it — but only the analyst's channel-split can rule out an acquisition shift.

Running it taught me

the parts the plan couldn't predict ✦
The best finding killed my own hypothesis.

"Extend the trial" was the obvious, PM-friendly answer — and it was wrong. Tom's interview retired it. The discipline that mattered was reporting that against my own and Rohan's prior, and making the non-recommendation as loud as the recommendations.

The quota saved the study.

Daniel and Sarah — the two low-fluency beginners I deliberately recruited — carried the entire overwhelm finding. Had I let the pool skew young and app-fluent like it wanted to, I'd have missed the strongest H1 evidence completely.

Synthesis is where the value is made.

The raw interviews gave me four themes. The work was realizing H2 and H4 weren't peers of H3 — one folds in, one collapses. Collapsing four findings into one sentence is what made the recommendation rank-able.

✦ the ending the plan set up
Three weeks bought Product a direction they could trust — and one expensive idea they were spared from building.

Whether the momentum arc moves D30 is for design to build and an A/B test to settle. Research's job was to make sure they build the right thing. On a simulated run, that's exactly where it landed.

D.
simulated case study · part two · 2026