Devon Park
AI Safety Researcher · Alignment, red teaming, interpretability
Published two workshop papers on jailbreak robustness. Runs an automated red-team suite with 1,200+ probes across three model families.
Artifact Gallery
Risk taxonomy with severity and likelihood for 30+ deployed systems.
Aug 20, 2026Instruction-hierarchy enforcement with indirect injection mitigation.
Jul 17, 20261,200+ adversarial probes across three model families with scoring.
Jun 14, 2026Capabilities, limitations, intended use, and red-team findings.
May 11, 2026Disparate-impact testing across demographics with visual report.
Apr 8, 2026Process Trace Summary
- Repository initializedJul 13
Created project scaffold with README and license.
- First commit pushedJul 15
Initial proof-of-concept with placeholder data.
- Evaluation harness addedAug 17
Wired up 50-question regression suite with LLM-as-judge.
- Peer review feedbackAug 19
Two reviewers flagged edge cases in retrieval fallback path.
- Iteration — fallback hardenedSep 21
Added retry + validation; eval score improved 12 points.
- Final submissionSep 23
Artifact submitted for oral defense scheduling.
Oral Defense Transcripts
Recorded Q&A from each verified oral defense session. Expand a session to read the transcript.
Walk us through the architecture of your artifact. Why did you choose this approach?
I chose a plan-and-execute topology because the task required multi-step retrieval with reflection. The plan node decomposes the query, sub-agents retrieve and draft in parallel, and a reflection node scores and routes for a second pass when below threshold.
What evaluation did you run, and what were the headline numbers?
I ran a 50-query regression suite scored by LLM-as-judge calibrated against a human panel (0.86 agreement). Baseline scored 71%; the reflection pass lifted it to 88% with a 14% latency cost, which stayed within budget.
Describe a failure mode you found and how you mitigated it.
Retrieval fallback returned stale context on schema changes. I added a freshness check + retry with a smaller context window, which reduced stale-grounded answers from 9% to under 2%.