Design tests

Design tests is an engineering skill installed by Flow. In this guide, you will learn:

  • The problem design tests solves
  • What the skill treats as observable, and why
  • How the skill runs and what it produces

Overview

design-tests makes the decisions that come before test code. It decides which behaviors earn a test, which layer each test runs at, what each test asserts, and where the fakes go. It works for every surface an AdonisJS app exposes: an HTTP request, a console command, a queued job, and an event listener.

Like every engineering skill, it answers one recurring design question and works on its own, outside the spec-driven workflow. It produces decisions, not test files. Writing the tests is left to you, your agent, or the build step.

The problem it solves

Left to its defaults, an AI agent writes the wrong tests. It writes a unit test for a controller and fakes the models the controller calls. It asserts that a service method ran instead of what the user sees. It bundles three effects into one test, and it pushes a form into a state the UI can never produce. Every one of those tests goes green, and none of them proves the feature works.

The result is a suite that is large and slow but still misses the bugs that matter. Those bugs live in the wiring between the router, middleware, validator, controller, and database, and a test that fakes the wiring never exercises it.

Design tests fixes this by making each test prove a behavior through the surface a real consumer uses, and by dropping any test that buys no confidence the suite does not already have.

What counts as observable

Every surface has a consumer, and a test asserts only what that consumer can see.

SurfaceIts consumerWhat a test asserts
An HTTP requestA person or clientWhere they land, what is rendered or returned, the status
A console commandWhoever ran itThe exit code and what was printed
A queued jobThe system that runs itThe state it left behind and the effects it fired
An event listenerThe system that runs itThe state it left behind and the effects it fired

Everything else is internal, such as a component instance, a service call, or the fact that a method ran. A test that asserts internals breaks when you refactor and passes when the feature is broken, so the skill never plans one.

How to use it

The skill reads your project's testing harness doc (testing.md) for the suites, helpers, and fakes it can plan against. That doc ships with the @adonisjs/core harness docs, and the skill stops if it is missing. It also reads the spec files that already cover the surfaces in scope, because the existing suite decides what counts as new confidence.

You can invoke it directly when you are deciding what to test for a feature, whether a unit test is warranted, or whether something should be mocked:

/flow-design-tests

The agent may also reach for it on its own when it recognizes that kind of question, and another skill can hand it a list of behaviors to turn into test decisions. See your agent's guide for the exact command syntax.

design-tests is a standalone skill, so it installs even when you skip the workflow skills. See Installation for how to select it.

How it decides

The skill proposes its decisions and has you correct them. It stops and asks whenever a choice would bind work beyond the current change. It works through four steps.

  • Place each behavior on a surface. Every behavior names the surface that exposes it and what its consumer would see. A behavior that no surface exposes is not testable as stated, so the skill hands it back to you.
  • Decide which behaviors earn a test, and at which layer. Every test defaults to the suite that drives the real surface, such as a functional HTTP test or a browser test. A unit test earns its place only when a piece of logic owns branching, a calculation, or a state transition that is awkward to reach through the surface. The skill settles each candidate with one question: if you deleted this unit test, what bug would slip through that no full-surface test catches? Behaviors already covered elsewhere, getters, configuration, and framework behavior are dropped, since types and linting cover them.
  • Pin what each test asserts. Each test proves one behavior, because a bundled failure does not say which effect broke. Tests reach only the states the surface can produce, and wait on conditions (an element visible, a row written) rather than on the clock.
  • Place the fakes. Fakes sit at the process boundary. That means the built-in fakes for mail, queues, and the like, or a container swap for your own service that wraps external IO such as a payment gateway or a clock.
Warning

The skill stops the run when a plan tries to unit test a controller, a command, or a job handler. It also stops when a test fakes your own service, model, or controller. Each of these is an orchestrator, and a unit test only checks its wiring against fakes.

Move the boundary outward instead. Test the controller through a real HTTP request, and fake the external API that your service wraps.

What it produces

Design tests produces decisions, not test code. It opens with the harness docs it read, followed by up to four blocks:

  • Tests list each kept behavior with its surface, its layer, and a title written as a behavior sentence.
  • Assertions state what each test asserts and what it deliberately leaves out.
  • Doubles name each fake or container swap, the test it serves, and what it stands in for.
  • Dropped lists every behavior that earned no test, with its reason.

The test lint rules enforce the structural side of the same discipline, such as keeping shared state and helpers out of spec files.

Next steps

  • Read Test lint rules to enforce test-suite conventions with ESLint.
  • Read Assert for test planning inside the spec-driven workflow.
Terms & License Agreement