Nextcraft
Liam Chen

Liam Chen

Agent Reliability Engineer · Observability, tracing, SRE

AI Orchestration EngineerVerified
Match Score90%

Owns tracing for an agent platform serving 200+ internal teams. Built a token-cost alerting system that saved $1.2M/year.

9Artifacts
86Defense Score (avg)
10Microcredentials
5Competencies Mastered

Process Trace Summary

  1. Repository initializedJul 19

    Created project scaffold with README and license.

  2. First commit pushedJul 21

    Initial proof-of-concept with placeholder data.

  3. Evaluation harness addedAug 23

    Wired up 50-question regression suite with LLM-as-judge.

  4. Peer review feedbackAug 25

    Two reviewers flagged edge cases in retrieval fallback path.

  5. Iteration — fallback hardenedSep 27

    Added retry + validation; eval score improved 12 points.

  6. Final submissionSep 29

    Artifact submitted for oral defense scheduling.

Oral Defense Transcripts

Recorded Q&A from each verified oral defense session. Expand a session to read the transcript.

Q1

Walk us through the architecture of your artifact. Why did you choose this approach?

A

I chose a plan-and-execute topology because the task required multi-step retrieval with reflection. The plan node decomposes the query, sub-agents retrieve and draft in parallel, and a reflection node scores and routes for a second pass when below threshold.

Q2

What evaluation did you run, and what were the headline numbers?

A

I ran a 50-query regression suite scored by LLM-as-judge calibrated against a human panel (0.86 agreement). Baseline scored 71%; the reflection pass lifted it to 88% with a 14% latency cost, which stayed within budget.

Q3

Describe a failure mode you found and how you mitigated it.

A

Retrieval fallback returned stale context on schema changes. I added a freshness check + retry with a smaller context window, which reduced stale-grounded answers from 9% to under 2%.

Competency Progress

Agent Architecture PatternsAvailable
Multi-Agent Communication Mastered
Tool Use & Function Calling Mastered
Prompt Engineering Fundamentals Mastered
RAG Pipeline Design Mastered
Vector Databases & Embeddings Mastered
LLM Evaluation & Metrics In Progress
Guardrails & Output Validation In Progress

Microcredential Verification

RAG Pipeline DesignIssued Jun 16, 2026
86Verified
Agent Architecture PatternsIssued Jul 16, 2026
87Verified
Multi-Agent CommunicationIssued Aug 16, 2026
88Verified
Tool Use & Function CallingIssued Sep 16, 2026
89Verified
Prompt Engineering FundamentalsIssued Oct 16, 2026
90Verified