Priya Iyer
Human-AI Product Designer · Agentic UX, trust calibration
Designed the transparency system for an assistant with 2M MAU. Portfolio includes a full human-in-the-loop review pattern library.
Artifact Gallery
User mental model alignment research with over-trust mitigations.
Aug 20, 2026Confidence indicators, source attribution, and model limitation disclosure.
Jul 17, 2026Figma prototype for delegating, interrupting, and reviewing autonomous agents.
Jun 14, 2026Multi-turn dialogue repair flows reducing escalations by 35%.
May 11, 2026Character design for AI assistants with contextual tone adaptation.
Apr 8, 2026Process Trace Summary
- Repository initializedJul 14
Created project scaffold with README and license.
- First commit pushedJul 16
Initial proof-of-concept with placeholder data.
- Evaluation harness addedAug 18
Wired up 50-question regression suite with LLM-as-judge.
- Peer review feedbackAug 20
Two reviewers flagged edge cases in retrieval fallback path.
- Iteration — fallback hardenedSep 22
Added retry + validation; eval score improved 12 points.
- Final submissionSep 24
Artifact submitted for oral defense scheduling.
Oral Defense Transcripts
Recorded Q&A from each verified oral defense session. Expand a session to read the transcript.
Walk us through the architecture of your artifact. Why did you choose this approach?
I chose a plan-and-execute topology because the task required multi-step retrieval with reflection. The plan node decomposes the query, sub-agents retrieve and draft in parallel, and a reflection node scores and routes for a second pass when below threshold.
What evaluation did you run, and what were the headline numbers?
I ran a 50-query regression suite scored by LLM-as-judge calibrated against a human panel (0.86 agreement). Baseline scored 71%; the reflection pass lifted it to 88% with a 14% latency cost, which stayed within budget.
Describe a failure mode you found and how you mitigated it.
Retrieval fallback returned stale context on schema changes. I added a freshness check + retry with a smaller context window, which reduced stale-grounded answers from 9% to under 2%.