AI Quality Engineer · LLM Evaluation & Test Automation

Satish Chilkaka

QA/SDET engineer with 9+ years of test automation experience, now specializing in AI quality and LLM evaluation. Built a production RAG chatbot with an automated LLM-as-judge evaluation harness. Deep expertise in Python, Playwright, Cypress, and CI/CD quality gates across regulated medical-device, fintech, and travel platforms.

Richmond Hill, ON · Open to remote & hybrid

About

I'm a QA/SDET engineer with 9+ years of test automation experience, now specializing in AI quality and LLM evaluation. I built a production RAG chatbot with an automated LLM-as-judge evaluation harness covering positive and negative-path test cases — applying systematic test design to the hardest new QA problem: verifying non-deterministic LLM output.

My background spans regulated environments (medical devices), fintech, automotive, and travel platforms, with deep expertise in Python, Playwright, Cypress, and CI/CD quality gates. I'm currently building LLM-powered applications and evaluation tooling.

Experience

May 2023 — Present Current

Senior QA Engineer

NeuroFlex — Toronto

  • Own end-to-end QA across web, Windows, and VR-based brain analysis applications — functional, regression, integration, and exploratory testing for clinician-facing workflows.
  • Built dual-stack automation: TypeScript + Cypress (v1) and Python + Playwright (v2), cutting manual regression effort ~40% per release.
  • Designed test plans, traceability matrices, and validation reports aligned with medical-device QA standards.
  • Integrated automated tests into Bitbucket Pipelines and Jenkins — per-PR execution, merges blocked on regression failure.
  • Stood up Grafana, Prometheus, and SonarCloud dashboards plus Locust performance testing for shared visibility into production stability.
Jan 2021 — Apr 2023

Lead QA Automation Engineer

Sherpa — Toronto

  • Architected the team's first BDD automation framework (Cypress, Playwright) — nightly regression from ~6 hours manual to a 45-minute automated run.
  • Led REST API testing across 80+ endpoints (auth, payments, submissions) with Postman, Cypress, and Axios.
  • Mentored two junior QA engineers; established code-review standards for test code; embedded suites in GitHub Actions and Jenkins with Slack failure alerts.
Jul 2016 — Dec 2020

Senior QA Analyst

AutoServe1 — Toronto

  • Increased automated coverage 25% with Node.js + Cypress suites for multi-role permission and invoicing workflows.
  • Ran load/stress testing with JMeter and BlazeMeter; owned functional, API, regression, and UAT planning across Agile sprints.
Earlier roles

Earlier QA Roles — SwaBz Systems · Mdreams (D3) · H-Line Soft Solutions — Toronto / Montreal / India (Dec 2011 — Jun 2016): manual, functional, regression, and API testing; defect tracking with JIRA and HP Quality Center.

Projects

RAG Chatbot & LLM Evaluation Harness

LLM + Evaluation

A production RAG-based chatbot — FastAPI backend with Groq (Llama-3.3-70B), hosted on Render, with a React frontend. Paired with an automated LLM-as-judge evaluation pipeline (Python, Gemini judge model) scoring actual vs. expected answers from a versioned test-case suite, including negative-path cases: hallucination checks, out-of-scope queries, prompt-injection attempts, and refusal correctness. Eval runs are treated as CI-style regression gates for every model/prompt change.

PythonFastAPIGroqGeminiReactLLM-as-judge

Spendly

Full-Stack Web App + AI

A full-stack personal finance application — a Django REST API backend and a React single-page frontend covering data modeling, authentication, and CRUD workflows end to end. Integrated backend-proxied LLM expense categorization (Gemini 2.0 Flash via OpenRouter) — prompt design, structured JSON output parsing, and failure-mode handling.

PythonDjangoReactGeminiOpenRouterREST

AI QA Agent In progress

LLM + Browser Automation

An autonomous agent that explores a web app, generates and runs tests via Playwright, and reports defects — with an evaluation harness that measures how many real bugs it catches. Built to run on local (Ollama) or cloud models.

PythonPlaywrightLLMAgentsEvaluation

Technical skills

AI & LLM engineering

LLM EvaluationLLM-as-judgeRAG PipelinesPrompt Engineering Negative-Path TestingOpenRouterGroqGeminiOpenAI-compatible APIsFastAPIGreat Expectations

Test automation

PlaywrightCypressSeleniumPytestJestPostmanREST Assured

Languages

PythonTypeScriptJavaScriptNode.jsSQL

CI/CD & DevOps

GitHub ActionsJenkinsBitbucket PipelinesDocker

Cloud & monitoring

AWSAzureGCPGrafanaPrometheusSonarCloudLocustJMeter

Databases

PostgreSQLMySQLSQL ServerMongoDBFirestoreOracle

Test management & methodologies

TestRailJiraConfluenceHP Quality Center Agile/ScrumBDDTDDShift-Left QARisk-Based Testing

Outside of my day job

Websites & web apps for small businesses

I design and build fast, modern sites and small web apps for local businesses — clean code, mobile-first layouts, and easy updates. If you need a presence that matches the quality of your work, I'd love to hear from you.

Discuss a project

Education & publications

Education

  • M.S., Software Engineering Concordia University, Montreal — 2015
  • B.E., Engineering JNTU University, India — 2011

Community & building with AI

Engineering community

I share practical engineering and testing ideas, encourage quality ownership, and help teams improve reliability through better tooling and coverage.

If you're building a developer community or running discussions on automation, APIs, or AI tooling, I'm happy to collaborate.

Building with AI

I'm currently building LLM-powered tools — agents and evaluation harnesses that bring engineering rigor to non-deterministic systems.

  • Web-testing agents that explore apps and generate tests
  • RAG systems with retrieval-quality evaluation
  • LLM-as-judge pipelines with calibration against human labels

Let's connect

Hiring managers and teams: I'm interested in software engineering roles — full-stack, backend, or SDET — where building well matters.