Fictional resume example. Names, employment histories and results are illustrative, not an actual employee record.
Lewis O'Connor
Senior AI Engineer with experience in retrieval quality, LLM evaluation, safe tool use, and inference cost. Practical work includes Python, RAG, LLM evaluation, Vector search.
Experience
NYU Langone
New York · United States
Senior AI Engineer
Jan 2021 - now
- As workstream lead, built a Python RAG service over permission-filtered support documents. Compared retrieval depth and prompt engineering variants using fixed answer and citation tests; selected the configuration that resolved untraceable citations without exposing restricted documents.
- With responsibility for the review standard, added automated testing to CI/CD for refusal behavior, citation validity, and prompt injection; connected monitoring to the same failure categories so a fluent but unsupported answer could block a release.
- As workstream lead, deployed the Python service as a versioned Docker image on AWS ECS, with scoped IAM access to source documents. Validated health checks and rollback against the previous image; released the service only after citation and permission tests passed.
- Built an evaluation set separating retrieval errors from unsupported answers, checked document permissions before retrieval, and recorded citation accuracy alongside latency and token cost.
- Owned a RAG evaluation using synthetic clinical questions and Python fixtures, resolved unsupported citations and completed human review before a limited demonstration.
- Built automated testing for prompt versions using Docker and CI/CD, reduced repeated citation failures by 34% on the held-out fixture and published observability checks.
Microsoft
Redmond, Washington · United States
AI Engineer
Mar 2016 - Dec 2020
- Added approval gates and idempotency keys to an agent tool workflow that updates customer records, with a replayable audit log; aligned review standards and decision owners across the participating teams. Working with partner teams, prevented a retried tool call from creating duplicate customer updates.
- Investigated a discrepancy between a published metric and its source records, traced the transformation that changed the population, and corrected the calculation with a reproducible query.
- Compared the last successful data refresh with a failed run, separated missing source data from transformation errors, and reran only the affected interval. Kept the original query and corrected result together for review.
Selected project
AI Engineer — independent case study
Cross-functional project lead
Mar 2023 - Jul 2023
- Added monitoring for token usage and p95 latency by request type, then routed simple requests to a smaller model with an explicit fallback; aligned review standards and decision owners across the participating teams
- Generated synthetic source records with duplicates, late arrivals and corrected values; wrote assertions for row counts and key uniqueness, and recorded the expected effect of each case on the reported metric.
- Compared the analytical output with a manually calculated reference table, traced differences to a transformation step, and retained a data dictionary and rerun instructions alongside the corrected query.
- Owned the synthetic-data validation using a manually calculated reference, resolved duplicate-key inflation and completed a notebook that reproduces the corrected totals.
Education
University of Washington
Seattle, Washington · United States
B.S. Computer Science
Sep 2008 - Jun 2012
Relevant coursework: Algorithms, operating systems, databases, computer networks
Skills
Role expertise
Python · RAG · LLM evaluation · Vector search · Observability
Publications
- Published an independent methods note using reproducible queries, explaining the data grain, excluded records and sensitivity of the result to a changed denominator.


