Bareerah Iftikhar
← All work/EdTech · AI trust · Design assessment

How I'd help teachers trust an AI recommendation

Alef's engine already knew what each student needed — but teachers accepted every suggestion blind or ignored them all. I reframed it from a model-accuracy problem to a trust problem: one flow where the reason arrives before the action.

/ In 30 seconds4 min full read
Problem
Alef's engine already knew what each student needed. Teachers accepted 61% of recommendations in under 8 seconds without opening the reason, and never opened 33% at all.
My call
Treated it as a trust problem, not a model problem: one spine — Notice → Understand → Decide → Assign → Verify — where the reason and its limits arrive before the action, and only risky picks are gated.
Proof
An assessment, not a shipped product. I set the metric, the guardrails, and the kill criteria I'd hold it to instead of quoting an uplift I never measured.
Client
Alef Education — Lead Product Designer exercise
Role
Lead Product Designer exercise · self-directed
Year
2026
Tools
Figma · Stitch · Claude · ChatGPT · Gemini · Paper
/ Key screens
  • Reason-first recommendations
    Reason-first recommendations
  • "See why" — evidence panel
    "See why" — evidence panel
  • Off-level guardrail
    Off-level guardrail
  • Assigned — confirmation
    Assigned — confirmation
  • Adjust or dismiss
    Adjust or dismiss
/ How I worked

Method, insight, decision, tradeoff.

The four beats that turn a project into evidence — how I researched it, what I learned, the call I made, and the compromise I accepted to make the call.

01Method

A close read of the brief and its pilot data, a conversation with a practising teacher, and AI-accelerated domain research (Claude, ChatGPT) plus parallel ideation passes (Gemini, Stitch) run against my own paper wireframes.

02Insight

Trust was failing in two directions at once: 61% of accepted recommendations were actioned in under 8 seconds without opening the rationale, while 33% were never opened at all. Acting costs nothing; understanding costs effort.

03Decision

Put the reason and its limits before the action, and gate only the risky picks. Off-level students stay out by default; safe assigns stay fast.

04Tradeoff

Excluding off-level students by default is safer but can hide a stretch a teacher intended. I'd A/B a lighter "show all" control rather than pretend the call is settled.

/ How I ship

Phases, decisions, artifacts, outcomes.

The actual shape of the work — not the marketing version. Each phase lists the calls I made, what shipped, and what moved.

01

Frame

human-led
Decisions
  • Read the brief for the real failure before designing: the model has residual risk, but the product failure is trust.
  • Treat the 12% audit rate and the one logged wrong-level assignment as design constraints, not footnotes.
Artifacts
  • Problem memo
  • Explicit assumption list
  • Scope: main flow, not every edge case
Outcome

Named the two opposite failures — over-trust and under-trust — and the shared root cause.

02

Explore

AI-assisted + human-led
Decisions
  • Use AI to map the EdTech domain fast instead of studying competitors one by one.
  • Look outside EdTech at how products present AI suggestions — that's where the confidence chip came from.
  • Speak to a real teacher about assigning work, so the design sits on a person, not a guess.
Artifacts
  • Cross-domain teardown of AI suggestion patterns
  • Paper wireframes alongside independent Claude / ChatGPT / Stitch drafts
  • One merged direction chosen by hand
Outcome

A three-tier confidence language (High / Medium / Check first) and a reason-first card.

03

Ship

AI + human
Decisions
  • The primary button says "Review & assign", never "Accept".
  • Off-level and low-signal students are unticked by default; adding one back takes a deliberate tick and a warning.
  • The see-why panel names what the system may not know, not only the evidence that supports the pick.
Artifacts
  • Core teacher flow: recommendations, see-why, adjust/dismiss, guardrail, assigned, assign-without-AI, impact
  • Three reusable primitives and one system rule
  • Stitch drafts rebuilt by hand in Figma
Outcome

One flow that carries every content type — lesson, video, assessment, activity, practice.

04

Hold

if built
Decisions
  • Measure informed action, not acceptance rate — acceptance rewards rushing.
  • Stop the rollout if wrong-level assignments rise, even if usage rises with them.
Artifacts
  • Instrumentation plan for See why / Review & assign / Adjust / Dismiss / off-level re-add
  • A/B against today's flow plus follow-up sessions with 6–8 teachers
Outcome

A hypothesis with a kill condition attached, rather than a success number I didn't earn.

/ 01 · What this is

What this is

A Lead Product Designer exercise for Alef Education, published in full. It is an assessment, not shipped work, and it is labelled that way everywhere it appears — the value here is the reasoning, not a result I can't claim.

The engine already existed and could suggest student–content matches with a confidence signal. Retraining it was not the job. Designing for suggestions that are usually right and occasionally wrong was.

/ 02 · The reframe: not only a model problem, a trust problem

The reframe: not only a model problem, a trust problem

An independent audit judged roughly 12% of recommendations questionable or wrong, and one wrong-level task reached a struggling reader before anyone noticed. That risk is real. But the interface did nothing to contain it, and it did nothing to earn belief either.

So teachers split two ways. 61% of accepted recommendations were actioned in under eight seconds without the rationale ever being opened. 33% were never opened at all, and 72% of dismissals happened without the reasoning being read. Both failures share one root cause: acting on a recommendation costs nothing, understanding one takes extra effort. More filters and bigger tables make that worse. The fix is making the informed choice the easy one.

61%
Over-trust — accepted in <8s, rationale unopened
33%
Under-trust — recommendations never opened
~12%
Recommendations judged questionable by audit
86%
Assignments still going to the whole class
/ 04 · How I worked this brief

How I worked this brief

Framing was mine: I read the brief closely to find the real problem before drawing anything. Domain research was AI-assisted — with limited time, I used Claude and ChatGPT to map EdTech products and their patterns quickly, then studied how products outside EdTech present AI suggestions. The three-tier confidence chip came from that detour, not from another EdTech dashboard.

Then a real teacher, on her real frustrations with assigning work. Ideation ran in parallel: paper wireframes by hand while Claude, ChatGPT, Gemini and Stitch produced independent drafts from the same brief. I compared every version and merged the strongest into one. Stitch produced rough high-level screens; the final craft and interaction decisions were made by hand in Figma.

The honest bottom line: AI made me faster, not smarter. The domain understanding, the trust decisions, and the craft are mine.

/ Today vs redesigned
Journey map of the current assignment flow with trust dipping at the decide stage
Today · where trust breaks
Journey map of the redesigned flow with a design response at each stage
Redesigned · trust holds

The same five-stage spine mapped twice. Today, trust dips at Decide — skim, guess, fall back to the whole class, never learn if it helped. Redesigned, each stage has a response that holds trust across the flow.

/ 06 · 1 · Reason first, confidence up front

1 · Reason first, confidence up front

The reason sits on the card itself, with a confidence level in words — High, Medium, or Check first. The primary button reads "Review & assign", never "Accept". The wording nudges a look before the action.

Why not a one-click Accept and a numeric score: Accept is what produced the 61% blind accepts, and a percentage invites false precision about a model that is wrong roughly one time in eight. The chip says what to do, not how sure a model claims to be.

Recommendation list screen showing reason-first cards with confidence chips
Reason-first cards with confidence and evidence chips. The third card flags an off-level match before it's even opened.
/ 08 · 2 · Show the working, and the limits

2 · Show the working, and the limits

One tap opens the evidence: the exact scores, plus an amber note naming what the system may not know. Showing the limits is what earns the right amount of trust rather than the maximum amount.

Why not show only the scores that argue for the pick: a reason that only defends itself breeds over-trust. Naming the limit is what keeps trust honest.

Explanation panel showing evidence scores and a limits note
The see-why panel: evidence table plus an explicit "what the system may not know" note.
/ 10 · 3 · Adjust and dismiss carry equal weight

3 · Adjust and dismiss carry equal weight

Adjust sits next to Assign, so editing who's included or what's assigned isn't buried. Dismiss asks for a one-tap reason, which turns a skip into feedback instead of silence.

Recommendation card with adjust and dismiss actions beside assign
Adjust and Dismiss placed beside Assign — equal weight, not hidden in an overflow menu.
/ 12 · 4 · Friction only where risk is higher

4 · Friction only where risk is higher

Before assigning, the teacher confirms who's in. On-level students are ticked by default; off-level and low-signal students are left out. Adding an off-level student takes a deliberate tick and a plain-language warning tied to the logged incident.

That's how the UX contains residual model risk without slowing every safe assign. Why not the alternatives: ticking off-level students by default is exactly what caused the logged mis-assignment, and hard-blocking them removes the teacher's judgment. Deliberate opt-in sits between the two.

Assignment guardrail screen warning that a student is off-level
"Ayesha is off-level for this task." A warning tied to a real incident, not a generic confirm dialog.
/ 14 · 5 · Who's in, who's held back, why

5 · Who's in, who's held back, why

The confirmation names who received the work, who was held back and why, and how long the assignment took — then offers impact or the next suggestion. Why not a silent success toast: naming who was held back keeps the teacher in the loop and sets up the verify step.

Assignment confirmation screen listing included and held-back students
Assignment summary: included, held back with a reason, and time to assign.
/ Three reusable primitives
Recommendation card pattern with reason, confidence and actions
Recommendation card
See why panel pattern showing evidence and known limits
See-why panel
Student row pattern with on-level and off-level badges
Student row with guardrail

The pattern isn't about fractions or Grade 6. The card, the panel, and the student row carry to any subject, grade, or future Alef product — the words change, the structure stays.

/ 18 · What I'd measure — and when I'd stop

What I'd measure — and when I'd stop

I didn't ship this, so I won't invent a success number. The primary metric is informed action: the teacher saw the reason, then acted. Blind accepts and never-opened both fall from 61% and 33%; wrong-level assignments fall; time to assign moves from 22 minutes toward under five.

Proof would come from instrumenting See why, Review & assign, Adjust, Dismiss and off-level re-adds, an A/B against today's flow, then sitting with six to eight teachers after a short assign task. Numbers show what they did; the conversation shows whether they understood.

The kill condition: if wrong-level assignments rise, I'd stop, even if usage rises with them. Faster is not better if students get the wrong work.

/ 19 · What I'd do differently, and what I don't know

What I'd do differently, and what I don't know

  • I'm working from the outside — Alef's team knows constraints I don't, and some choices will bump into them.
  • The whole design rests on the first recommendation and its explanation being genuinely good. The next pass is writing and testing that copy, not more screens.
  • The guardrail default is a judgment call: excluding off-level students is safer, but may hide a deliberate stretch. I'd A/B a lighter "show all" control.
  • The admin verify view is under-designed — enough to close the loop, not enough to prove intervention effectiveness.
  • Before any rollout: loading and error states, validating the three-tier confidence labels with real teachers, and instrumenting informed action first.
/ Outcomes

What shipped.

61% → target
blind accepts, the number the design attacks
1 spine
Notice → Understand → Decide → Assign → Verify
3 primitives
recommendation card, see-why panel, student guardrail
No invented uplift
measurement plan and a kill condition instead