Rebuilding Allie:

Allen's AI-powered

doubt solver

Rebuilding Allie: Allen's AI powered

doubt solver

Rebuilding Allie:

Allen's AI powered

doubt solver

Rebuilding Allie:

Allen's AI powered

doubt solver

How I rebuilt Allen's AI doubt solver to recover from a 35% decline after ChatGPT Study Mode and Gemini reshaped what students expect from AI.

How I rebuilt Allen's AI doubt solver to recover from a 35% decline after ChatGPT Study Mode and Gemini reshaped what students expect from AI.

How I rebuilt Allen's AI doubt solver to recover from a 35% decline

after ChatGPT Study Mode and Gemini reshaped what students expect from AI.

How I rebuilt Allen's AI doubt solver to recover from a 35% decline

after ChatGPT Study Mode and Gemini reshaped what students expect

from AI.

How I rebuilt Allen's AI doubt solver to recover from a 35% decline

after ChatGPT Study Mode and Gemini reshaped what students expect

from AI.

In eight weeks, Allie lost a third of its users, causing monthly doubts to plunge 35%. Not through a product failure but through a market reset we weren't prepared for. Students didn't use Allie less. They left for better products entirely. As the sole product designer, I redesigned the experience around four product bets that grew weekly doubts 120% and engaged users 79%.

Role · Sole Product Designer

Partners · Product Manager, Data Science, Engineering

80K → 176K +120%

weekly doubts

16K → 29K +79%

weekly engaged users

+23%

doubts per user

Rebuilding Allie:

Allen's AI-powered

doubt solver

How I rebuilt Allen's AI doubt solver to recover from a 35% decline after ChatGPT Study Mode and Gemini reshaped what students expect from AI.

The Market Reset

Two launches, six weeks, a third of our users gone

Late July 2025

ChatGPT Study Mode

launches free

August 2025

Gemini Pro free for students, one full year

October 2025

−35%

monthly doubts on Allie

ChatGPT Study Mode (late July):

A free Socratic tutor leading questions, scaffolding, hints instead of dumped answers. Free for everyone.

Gemini Pro free for students (August)

Google's flagship paid tier, free for one full year with student verification.

Within weeks, every student in India had a frontier AI tutor on their phone, for free. Overnight, Allie's one-shot text answer, one length, one language, one method became the weakest tutoring experience students had access to.


By October, monthly doubts had fallen 35%.

The Diagnosis

Students weren't using Allie less, they were leaving

Before designing anything, I needed to know what kind of problem this was: were students asking fewer doubts, or leaving entirely?

Decomposing the decline (Aug → Oct 2025)

Doubts per user

−7%

Users

−30%

Students who stayed behaved almost exactly as before. The decline was churn.

Students leaving meant they were actively choosing something better, so the question shifted from "how do we remind students Allie exists?" to "why did they stop trusting it?"

Why They Left

We surveyed 149 churned students over WhatsApp, ran four in-depth interviews, and benchmarked Allie's answers against the products students defected to.

The survey: ranking the complaints

"Correct but incomprehensible" the #1 churn reason

149 churned students, July's power users who had gone quiet, and irregular users who had left

entirely ranked their reasons for leaving:

1

AI text solutions were correct but difficult to understand: the highest-cited reason. Formatting, structure, and pedagogy were the issue, not accuracy.

2

AI text responses felt slow: first-response latency was 15 to 30s.

3

Video solutions were not available for many questions: coverage had degraded.

4

AI text solutions were incorrect: a smaller but non-trivial cohort.

5

Teacher solutions were not helpful: a long tail of quality issues.

Asked where they went instead:

ChatGPT

Gemini

Google AI Overviews

What interviews revealed

The interviews put a mechanism behind the survey. Despite spanning very different profiles:

  1. Power user who defected to ChatGPT because Allie ignored custom formatting instructions and failed to spot errors.

  2. Speed purist resolving 90% of doubts via Google due to high latency (15 to 30s ) and messy image replies.

  3. Below-average student who returned to offline help due to encountering lengthy, inaccurate answers and zero Hinglish support.

Every interview pointed to the same pattern:

Trust broke after ~2 poor responses

Try Allie, get a bad answer, try once more, never return. Three of four students described it unprompted.

Understanding, not just Answers

Students were rarely asking for the final answer, they wanted

the approach, the concept, the mistake.

AI needs to teach like a teacher

Implicitly stated across multiple interviews. Students disliked Allie partly because the method differed from what their teacher taught.

"Approach nahi samjha, concept nahi

samjha.

Tell me what mistake I made."

Recurring phrases across churned-student interviews

What students did beyond the research

Apart from the research we also noticed a recurring pattern in the product: 70% of doubts originated from images of MCQs immediately after a test, yet asking one required leaving review and re-uploading the question. We hypothesised that doubts weren't disappearing but they were dying in between, and validated this through student interviews.

Doubt is born

Doubt is born

in test review

in test review

Leave review

Leave review

context left behind

context left behind

Open Doubts

Open Doubts

a separate destination

a separate destination

Re-upload Qs

Re-upload Qs

from scratch

from scratch

Doubt asked

Doubt asked

the few that survive

the few that survive

The no. of circles represents no. of doubts. Each step loses some doubts

Benchmark: Allie fails thinly while ChatGPT fails confidently

We tested Allie, ChatGPT, and Filo on the same 26 JEE & NEET doubts across three categories that map to how students actually use these tools. The benchmark showed where each solver wins and fails.

Allie failed quietly under-explanation that hid its decent core accuracy. ChatGPT failed

confidently with beautifully formatted answers that were factually wrong. Filo consistently delivered the correct, structured responses.

Benchmark 2: Defining a Quality Standard

This benchmark is the visual evidence: same 10+ 6th to 10th doubts submitted to Allie, Gemini, and GPAI side by side, judged on correctness, scannability, pedagogical depth, and visual aids.

The competitive gap wasn't just accuracy, it was presentation. Gemini and GPAI made the same information visually navigable. Allie buried it.

The Solution

Four bets shipped in sequence

Rather than one large rebuild, the solution was structured, each addressing a distinct churn driver and each independently shippable. The bets were sequenced to ship value early and use real-world adoption data to inform the harder pedagogical work that followed.

Churn driver

Bet

Hypothesis

Friction — doubts born in tests,

solved elsewhere

1· Solve doubts where they happen

Embed Allie in test review

Removing context re-entry unlocks lost

doubt volume

Rigidity — one format, one

language

2· Let students choose how they're taught

Response modes

Students who choose how they're

taught stop defecting

Quality — correct but

incomprehensible

3· Improve the structure of the response

Benchmark-derived spec

Fixing comprehension, not just

accuracy, rebuilds trust

Dead ends — slow, unhelpful teacher replies

4· Send fewer, better doubts to teachers

Friction + honest wait times

Fewer, higher-intent escalations let teachers keep pace

Bet 1: Solve doubts where they happen

Bet 1 put the context-switch hypothesis to the test: embed Allie inside review, pass the question context automatically, and the lost doubts should come back, which would lead to:

Drive engagement

Increase doubt-clearing volume from existing users who previously abandoned the high-friction process.

Lower the barrier to entry

Activate new users who found the standalone journey too tedious to attempt.

Thus, we planned on creating an integrated flow for Allie in the test/practice review such that context can be passed and user can ask doubts regarding the question seamlessly to Allie.

What should the entry point be?

First discussion was around where should the entry point for the contextual doubt be visible to the students. Should it be floating or natively build inside the footer? Both had their pros and trade-offs which were:

Native Footer

Covers nothing, only works when footer exists.

Floater

Easy to spot, scalable but covers content.

We went with the Floater as reach mattered more than a clean screen.

Doubts are born on many screens, not just test review, and the native version would have needed a separate build for each one. The floater gave us a single entry point we could put anywhere, and it was much easier for students to spot.

The cost: floater can cover part of solution. We placed it where content is least dense, but it still blocks visibility. Making it draggable or shrink on scroll is the obvious next step.

Initial Concept:

Clicking on the floater opens up a composer with chips which leads to the bottom-sheet.

After lots of discussion regarding feasibility and tradeoffs, we ditched the input-field approach due to:


  1. Chips felt buried next to the composer, making them feel like an afterthought rather than one of the primary ways to get started.

  2. It boxed us in. No room to grow into richer intents or surface more than a couple of chips.

  3. Extra dev effort to build this which would hamper the timeline.

Solution:

FTUX tooltip to educate the user.

Bottom sheet over a full-screen takeover so the question context remained visible above

The three quick action pills surface the most common intents at the first interaction, removing the need to type anything to get started.

Students could either start with common intents using quick actions or ask custom questions in their own words.

The bottom sheet can be expanded when students want to focus on longer explanations.

Alternative actions remain visible after the first response, allowing students to explore a different explanation path.

Outcome post launch:

80K → 176K +120%

weekly doubts

16K → 29K +80%

weekly engaged users

+23%

doubts per user

~60%

doubts originate here

The doubts were always there but

The flow was losing them.

Bet 2: Let students choose how they're taught

No single answer format could serve every intent. The same student wants a ten-second confirmation on Tuesday and a full walkthrough on Thursday and the power users we lost had said it directly: they left because they couldn't shape the response. So we gave them a control surface.

Detailed:

The existing one-shot full solution. Default for students who want the complete approach explained.

Guided:

A step-by-step Socratic walkthrough. Rather than giving the answer away, it acts as a tutor that guides the student to the solution.

Short:

A concise version of the same solution. For quick checks, validation, or when the student just needs confirmation.

Variations of how modes should be visible on the homepage

Pill embedded in the composer, cleanest approach but lack of visibility

Tab placed on top, copy of the subtext would change based on the mode

Similar to the previous version but placed below Allie

Tab carousel with subtext for each above the composer. Full visibility

Tabs embedded in the composer itself. based on the mode selected, the composer text changes.

We went with the Pill Embedded approach as we wanted the least amount of friction and decision fatigue when students ask doubt.

Detailed already works for many students, so a loud mode picker adds a choice for the few who want something different. The pill keeps the default path simple: type and send, while keeping the control accessible. Modes appear again after the first response, when most students decide.

The cost is visibility: this is the least noticeable of the five.

Solution:

Students decide how they want help before asking a doubt.

Modes bottom-sheet with subtext to help users choose confidently without trial and error.

Answer generated in Detailed mode.

Students typically decide whether an explanation is useful after the first response thus alternative modes appear after the first response.

In Guided mode, Allie asks progressive questions to help students reason through doubts.

Outcome post launch:

~10%

adoption of guided

~8%

adoption of short

~90%

satisfaction rate for both

Roughly one in five doubts uses a non-default mode by choice. And 90% satisfaction rate for both meant these modes were helping students solve doubts the way they preferred.

The web view of the updated doubts flow optimised for desktop:

Bet 3: Improve the structure of the response

Based on the benchmarks, I created guidelines for the response structure to fix the issues and improve the quality.

Step 1

Benchmark

36+ doubts scored across

five solvers, on correctness,

scannability, pedagogy,

visual aids.

Step 2

Failure taxonomy

Recurring, nameable

failures: buried answers, no

step bifurcation, no visuals,

method drift.

Step 3

Quality guidelines

Structured with headers, pointers, bifurcation, scannable in seconds, solved the way an Allen teacher would solve it.

Step 4

Shipped improvement

The Data Science team

optimized model output

against the guidelines.

"Quality" stopped being an opinion and became a checklist, one that mapped directly to why students actually churned. The clearest way to see it is in the responses themselves:

Question

Old Response

Wrong answer

No structure or bifurcation in the solution

Difficult to scan

New Response

Right answer

Properly structured with headers

Easy to scan

Bet 4: Send fewer, better doubts to teachers

Teacher answer quality ranked #5 among churn reasons and also hurting our NPS. With answers taking 24+ hours, students were left waiting. Since quality wasn't a product problem, I didn't redesign the destination but changed who reached it and what they expected.

Reduced prominence: "Ask teacher" shifted from a prominent quick action to a secondary option pill available which no longer looks like the default next step.

Before

After

Deliberate friction on escalation: Instead of one-tap escalation, students now add a mandatory note before sending a doubt. This filters low-intent escalations and gives teachers the context they need.

An honest wait time, upfront: The expected teacher response time is shown before escalation, making the wait an informed choice, not a surprise.

Bottom-sheet pops up with a mandatory note input to proceed and expected TAT.

Notes comes up as a chat bubble for the teacher to read.

Outcome post launch:

4% → 2%

Doubts assigned to teacher

27h → 15h

Teacher turnaround time

What still needs fixing

Latency 15-30s

The #2 churn reason, owned by engineering and data science, some answers are already pre-cached to reduce latency but we still need to find a workaround for this.

Video coverage

The #3 churn reason, questions without videos still dead-end. AI-generated explainers for popular questions are in progress; I'm designing that experience and validating it through usability testing.

Language Selection

Language was one of the important feedbacks that we had gained during our research. We were about to launch it but due to some technical issues we had to pause it for now.

Guided Mode Issue

Currently the guided mode only works for problem solving questions, for conceptual questions it loses its socratic feature. I have created a flow to raise and fix this concern but its still TBD.

Key Takeaways

Three takeaways that travel beyond this specific feature:

Diagnose before you design

A 35% engagement drop looked like a generic UX problem. Breaking it down into user churn vs. usage frequency revealed the real issue: students were leaving for better AI alternatives.

External benchmarks reset the bar, fast

ChatGPT Study Mode and Gemini raised user expectations overnight. The lesson wasn't to copy competitors. It was to continuously benchmark against evolving AI experiences, not just direct competitors.

Layered solutioning beats one big swing

Instead of rebuilding everything at once, we shipped four independent bets. Early wins validated the strategy with real user data before investing in more complex improvements.

That’s a wrap!

Please view in

desktop for now

Please view in

desktop for now

Create a free website with Framer, the website builder loved by startups, designers and agencies.