Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Josephine Murphy
Junior AI Engineer with experience in retrieval quality, LLM evaluation, safe tool use, and inference cost. Practical work includes Python, RAG, LLM evaluation, Vector search.
Experience
NYU Langone
New York · United States
Junior AI Engineer
Jul 2025 - now
- With a senior colleague reviewing the change, built a Python RAG service over permission-filtered support documents. Compared retrieval depth and prompt engineering variants using fixed answer and citation tests; selected the configuration that resolved untraceable citations without exposing restricted documents.
- Within an assigned workstream, added automated testing to CI/CD for refusal behavior, citation validity, and prompt injection; connected monitoring to the same failure categories so a fluent but unsupported answer could block a release.
- With a senior colleague reviewing the change, deployed the Python service as a versioned Docker image on AWS ECS, with scoped IAM access to source documents. Validated health checks and rollback against the previous image; released the service only after citation and permission tests passed.
- Within the assigned scope, built an evaluation set separating retrieval errors from unsupported answers, checked document permissions before retrieval, and recorded citation accuracy alongside latency and token cost.
- Within the assigned scope, owned a RAG evaluation using synthetic clinical questions and Python fixtures, resolved unsupported citations and completed human review before a limited demonstration.
- Within the assigned scope, built automated testing for prompt versions using Docker and CI/CD, reduced repeated citation failures by 34% on the held-out fixture and published observability checks.
Microsoft
Redmond, Washington · United States
AI Engineer Project Contributor
Mar 2024 - Jun 2025
- The team added approval gates and idempotency keys to an agent tool workflow that updates customer records, with a replayable audit log; my scope was data preparation, quality checks, and follow-up documentation under senior review. The team prevented a retried tool call from creating duplicate customer updates; documented my assigned work separately from the team result and confirmed it with a senior colleague.
- Within the assigned scope, investigated a discrepancy between a published metric and its source records, traced the transformation that changed the population, and corrected the calculation with a reproducible query.
- Within the assigned scope, compared the last successful data refresh with a failed run, separated missing source data from transformation errors, and reran only the affected interval. Kept the original query and corrected result together for review.
Selected project
AI Engineer — independent case study
Project team member
Sep 2024 - Feb 2025
- Within the assigned scope, the team added monitoring for token usage and p95 latency by request type, then routed simple requests to a smaller model with an explicit fallback; my scope was data preparation, quality checks, and follow-up documentation under senior review
- Generated synthetic source records with duplicates, late arrivals and corrected values; wrote assertions for row counts and key uniqueness, and recorded the expected effect of each case on the reported metric.
- Compared the analytical output with a manually calculated reference table, traced differences to a transformation step, and retained a data dictionary and rerun instructions alongside the corrected query.
- Owned the synthetic-data validation using a manually calculated reference, resolved duplicate-key inflation and completed a notebook that reproduces the corrected totals.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science
Sep 2021 - Jun 2025
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
Python · RAG · LLM evaluation · Vector search · Observability
Publications
- Published an independent methods note using reproducible queries, explaining the data grain, excluded records and sensitivity of the result to a changed denominator.


