QA test generation with AI: covering the paths nobody wrote tests for
An agent reads a change, its ticket, and the bugs that have bitten this code before, then writes tests that pin down the behavior that matters. Every generated test has to fail against a deliberately broken version of the code before it is allowed into the suite.
QA lead or the engineering manager who owns test strategy
A feature branch is ready for review, or a module is flagged as under-tested
01the problem and who owns it
Test suites grow where writing tests is easy, not where failures are expensive. The billing edge cases, the timezone math, and the checkout flow on a slow network are exactly the parts that ship with a happy-path test or none at all, because writing the hard tests takes longer than writing the feature.
QA owns the strategy and the release sign-off, but a small QA team cannot keep pace with every branch. Engineers are measured on shipping features. The result is a coverage number that looks fine and a suite that rarely catches the regression that actually reaches customers.
02what the AI does, step by step
- Read the change and its intentThe agent pulls the diff, the acceptance criteria from the ticket, and any linked design notes. Tests are written against what the feature is supposed to do, not merely against what the code currently does.
- Mine the bug historyClosed bugs and incident reports that touched the same files become test ideas. A regression that happened once on a module is the most likely one to happen again.
- Write tests in your existing styleUsing the framework, fixtures, factories, and helpers already in the repository, such as Jest, pytest, JUnit, or Playwright, the agent writes focused tests. It does not introduce a new framework or mock everything in sight.
- Run them and fix the obviousThe agent runs the new tests in a sandbox. Tests that fail because the test is wrong get repaired. Tests that fail because the code is wrong are reported as possible bugs, not edited until they pass.
- Prove each test can failA mutation step changes the code under test, such as flipping a condition or removing a check, and confirms each new test goes red. Tests that pass against broken code are deleted.
- Open a reviewable pull requestSurviving tests arrive as a separate pull request or commit, grouped by behavior, with a note on what each one protects and which suspected bugs it surfaced.
03systems it connects to
- Source control and CI. GitHub or GitLab with CI runners that execute tests in isolated containers.
- Test frameworks. Whatever the project already uses: Jest or Vitest, pytest, JUnit, Playwright or Cypress for browser flows.
- Issue tracker. Jira or Linear for acceptance criteria and bug history.
- Mutation tooling. Tools such as Stryker, PIT, or mutmut, or a lighter custom mutation step for the changed functions.
04human checkpoints
- Engineer review of every test. Generated tests merge only after the feature's author or a reviewer reads them. A test nobody understands becomes a test nobody maintains.
- Suspected bug triage. When the agent reports that code contradicts the acceptance criteria, a person decides whether the code or the criteria are wrong.
- QA sign-off on end-to-end flows. Browser and device tests for critical journeys are approved by QA, who own what counts as a release blocker.
05what to measure
- Mutation score on changed code. A better signal than line coverage of whether tests detect real faults.
- Regressions caught before merge. Failures in generated tests that traced to genuine bugs.
- Flaky test rate. Generated tests that fail intermittently. These must stay near zero or the suite loses trust.
- Suite runtime. Total CI time, so added tests do not quietly slow every merge.
06risks and guardrails
- Tests that enshrine bugs. A model writing tests from existing code will happily assert current wrong behavior. Grounding tests in acceptance criteria and keeping the bug triage checkpoint is the countermeasure.
- Real data in fixtures. Never let the agent copy production records into fixtures. Use factories and synthetic data, and scan generated files for anything that looks like PII or credentials.
- Sandbox escape. Tests run code. Execute them in containers without production credentials or network access to internal systems.
07build vs buy
IDE assistants already draft unit tests on request, and for individual engineers that is often enough. Commercial test-generation products exist for specific languages and frameworks and are worth a trial.
A custom pipeline earns its place when you want tests generated systematically on every branch, tied to your bug history and acceptance criteria, and filtered by mutation testing so only tests that catch faults survive.
08related playbooks
Browse every engineering playbook or the full library.
want this running in your business?
We can point a test-generation agent at your most fragile module, run it with mutation checks in CI, and show you which tests it wrote that would have caught last quarter's bugs.
See how we deliver it: ai coding orchestration.
book a call drop your number