
How I'd help teachers trust an AI recommendation
Alef's engine already knew what each student needed — but teachers accepted every suggestion blind or ignored them all. I reframed it from a model-accuracy problem to a trust problem: one flow where the reason arrives before the action.
- Problem
- Alef's engine already knew what each student needed. Teachers accepted 61% of recommendations in under 8 seconds without opening the reason, and never opened 33% at all.
- My call
- Treated it as a trust problem, not a model problem: one spine — Notice → Understand → Decide → Assign → Verify — where the reason and its limits arrive before the action, and only risky picks are gated.
- Proof
- An assessment, not a shipped product. I set the metric, the guardrails, and the kill criteria I'd hold it to instead of quoting an uplift I never measured.
- Client
- Alef Education — Lead Product Designer exercise
- Role
- Lead Product Designer exercise · self-directed
- Year
- 2026
- Tools
- Figma · Stitch · Claude · ChatGPT · Gemini · Paper
Method, insight, decision, tradeoff.
The four beats that turn a project into evidence — how I researched it, what I learned, the call I made, and the compromise I accepted to make the call.
A close read of the brief and its pilot data, a conversation with a practising teacher, and AI-accelerated domain research (Claude, ChatGPT) plus parallel ideation passes (Gemini, Stitch) run against my own paper wireframes.
Trust was failing in two directions at once: 61% of accepted recommendations were actioned in under 8 seconds without opening the rationale, while 33% were never opened at all. Acting costs nothing; understanding costs effort.
Put the reason and its limits before the action, and gate only the risky picks. Off-level students stay out by default; safe assigns stay fast.
Excluding off-level students by default is safer but can hide a stretch a teacher intended. I'd A/B a lighter "show all" control rather than pretend the call is settled.
Phases, decisions, artifacts, outcomes.
The actual shape of the work — not the marketing version. Each phase lists the calls I made, what shipped, and what moved.
Frame
- Read the brief for the real failure before designing: the model has residual risk, but the product failure is trust.
- Treat the 12% audit rate and the one logged wrong-level assignment as design constraints, not footnotes.
- Problem memo
- Explicit assumption list
- Scope: main flow, not every edge case
Named the two opposite failures — over-trust and under-trust — and the shared root cause.
Explore
- Use AI to map the EdTech domain fast instead of studying competitors one by one.
- Look outside EdTech at how products present AI suggestions — that's where the confidence chip came from.
- Speak to a real teacher about assigning work, so the design sits on a person, not a guess.
- Cross-domain teardown of AI suggestion patterns
- Paper wireframes alongside independent Claude / ChatGPT / Stitch drafts
- One merged direction chosen by hand
A three-tier confidence language (High / Medium / Check first) and a reason-first card.
Ship
- The primary button says "Review & assign", never "Accept".
- Off-level and low-signal students are unticked by default; adding one back takes a deliberate tick and a warning.
- The see-why panel names what the system may not know, not only the evidence that supports the pick.
- Core teacher flow: recommendations, see-why, adjust/dismiss, guardrail, assigned, assign-without-AI, impact
- Three reusable primitives and one system rule
- Stitch drafts rebuilt by hand in Figma
One flow that carries every content type — lesson, video, assessment, activity, practice.
Hold
- Measure informed action, not acceptance rate — acceptance rewards rushing.
- Stop the rollout if wrong-level assignments rise, even if usage rises with them.
- Instrumentation plan for See why / Review & assign / Adjust / Dismiss / off-level re-add
- A/B against today's flow plus follow-up sessions with 6–8 teachers
A hypothesis with a kill condition attached, rather than a success number I didn't earn.
What this is
A Lead Product Designer exercise for Alef Education, published in full. It is an assessment, not shipped work, and it is labelled that way everywhere it appears — the value here is the reasoning, not a result I can't claim.
The engine already existed and could suggest student–content matches with a confidence signal. Retraining it was not the job. Designing for suggestions that are usually right and occasionally wrong was.
The reframe: not only a model problem, a trust problem
An independent audit judged roughly 12% of recommendations questionable or wrong, and one wrong-level task reached a struggling reader before anyone noticed. That risk is real. But the interface did nothing to contain it, and it did nothing to earn belief either.
So teachers split two ways. 61% of accepted recommendations were actioned in under eight seconds without the rationale ever being opened. 33% were never opened at all, and 72% of dismissals happened without the reasoning being read. Both failures share one root cause: acting on a recommendation costs nothing, understanding one takes extra effort. More filters and bigger tables make that worse. The fix is making the informed choice the easy one.
How I worked this brief
Framing was mine: I read the brief closely to find the real problem before drawing anything. Domain research was AI-assisted — with limited time, I used Claude and ChatGPT to map EdTech products and their patterns quickly, then studied how products outside EdTech present AI suggestions. The three-tier confidence chip came from that detour, not from another EdTech dashboard.
Then a real teacher, on her real frustrations with assigning work. Ideation ran in parallel: paper wireframes by hand while Claude, ChatGPT, Gemini and Stitch produced independent drafts from the same brief. I compared every version and merged the strongest into one. Stitch produced rough high-level screens; the final craft and interaction decisions were made by hand in Figma.
The honest bottom line: AI made me faster, not smarter. The domain understanding, the trust decisions, and the craft are mine.


The same five-stage spine mapped twice. Today, trust dips at Decide — skim, guess, fall back to the whole class, never learn if it helped. Redesigned, each stage has a response that holds trust across the flow.
1 · Reason first, confidence up front
The reason sits on the card itself, with a confidence level in words — High, Medium, or Check first. The primary button reads "Review & assign", never "Accept". The wording nudges a look before the action.
Why not a one-click Accept and a numeric score: Accept is what produced the 61% blind accepts, and a percentage invites false precision about a model that is wrong roughly one time in eight. The chip says what to do, not how sure a model claims to be.

2 · Show the working, and the limits
One tap opens the evidence: the exact scores, plus an amber note naming what the system may not know. Showing the limits is what earns the right amount of trust rather than the maximum amount.
Why not show only the scores that argue for the pick: a reason that only defends itself breeds over-trust. Naming the limit is what keeps trust honest.

3 · Adjust and dismiss carry equal weight
Adjust sits next to Assign, so editing who's included or what's assigned isn't buried. Dismiss asks for a one-tap reason, which turns a skip into feedback instead of silence.

4 · Friction only where risk is higher
Before assigning, the teacher confirms who's in. On-level students are ticked by default; off-level and low-signal students are left out. Adding an off-level student takes a deliberate tick and a plain-language warning tied to the logged incident.
That's how the UX contains residual model risk without slowing every safe assign. Why not the alternatives: ticking off-level students by default is exactly what caused the logged mis-assignment, and hard-blocking them removes the teacher's judgment. Deliberate opt-in sits between the two.

5 · Who's in, who's held back, why
The confirmation names who received the work, who was held back and why, and how long the assignment took — then offers impact or the next suggestion. Why not a silent success toast: naming who was held back keeps the teacher in the loop and sets up the verify step.




The pattern isn't about fractions or Grade 6. The card, the panel, and the student row carry to any subject, grade, or future Alef product — the words change, the structure stays.
What I'd measure — and when I'd stop
I didn't ship this, so I won't invent a success number. The primary metric is informed action: the teacher saw the reason, then acted. Blind accepts and never-opened both fall from 61% and 33%; wrong-level assignments fall; time to assign moves from 22 minutes toward under five.
Proof would come from instrumenting See why, Review & assign, Adjust, Dismiss and off-level re-adds, an A/B against today's flow, then sitting with six to eight teachers after a short assign task. Numbers show what they did; the conversation shows whether they understood.
The kill condition: if wrong-level assignments rise, I'd stop, even if usage rises with them. Faster is not better if students get the wrong work.
What I'd do differently, and what I don't know
- I'm working from the outside — Alef's team knows constraints I don't, and some choices will bump into them.
- The whole design rests on the first recommendation and its explanation being genuinely good. The next pass is writing and testing that copy, not more screens.
- The guardrail default is a judgment call: excluding off-level students is safer, but may hide a deliberate stretch. I'd A/B a lighter "show all" control.
- The admin verify view is under-designed — enough to close the loop, not enough to prove intervention effectiveness.
- Before any rollout: loading and error states, validating the three-tier confidence labels with real teachers, and instrumenting informed action first.