Nextcraft
Devon Park

Devon Park

AI Safety Researcher · Alignment, red teaming, interpretability

AI Safety & Governance LeadVerified
Match Score91%

Published two workshop papers on jailbreak robustness. Runs an automated red-team suite with 1,200+ probes across three model families.

9Artifacts
90Defense Score (avg)
10Microcredentials
3Competencies Mastered

Process Trace Summary

  1. Repository initializedJul 13

    Created project scaffold with README and license.

  2. First commit pushedJul 15

    Initial proof-of-concept with placeholder data.

  3. Evaluation harness addedAug 17

    Wired up 50-question regression suite with LLM-as-judge.

  4. Peer review feedbackAug 19

    Two reviewers flagged edge cases in retrieval fallback path.

  5. Iteration — fallback hardenedSep 21

    Added retry + validation; eval score improved 12 points.

  6. Final submissionSep 23

    Artifact submitted for oral defense scheduling.

Oral Defense Transcripts

Recorded Q&A from each verified oral defense session. Expand a session to read the transcript.

Q1

Walk us through the architecture of your artifact. Why did you choose this approach?

A

I chose a plan-and-execute topology because the task required multi-step retrieval with reflection. The plan node decomposes the query, sub-agents retrieve and draft in parallel, and a reflection node scores and routes for a second pass when below threshold.

Q2

What evaluation did you run, and what were the headline numbers?

A

I ran a 50-query regression suite scored by LLM-as-judge calibrated against a human panel (0.86 agreement). Baseline scored 71%; the reflection pass lifted it to 88% with a 14% latency cost, which stayed within budget.

Q3

Describe a failure mode you found and how you mitigated it.

A

Retrieval fallback returned stale context on schema changes. I added a freshness check + retry with a smaller context window, which reduced stale-grounded answers from 9% to under 2%.

Competency Progress

Alignment Fundamentals Mastered
Red Teaming Methodologies Mastered
Model Card Authoring In Progress
Bias Auditing In Progress
AI Policy Frameworks In Progress
Risk Taxonomy & ClassificationAvailable
Interpretability TechniquesAvailable
Incident Response for AI Mastered

Microcredential Verification

Alignment FundamentalsIssued Jun 10, 2026
91Verified
Model Card AuthoringIssued Jul 10, 2026
92Verified
Bias AuditingIssued Aug 10, 2026
93Verified