self_test 2.3.0 copy "self_test: ^2.3.0" to clipboard
self_test: ^2.3.0 copied to clipboard

Automated regression testing for Flutter through direct callback invocation. Run tests in live app environments without UI simulation.

2.3.0 - 2026-09-22 #

Three changes, all of them written because something upstream of a device run could be wrong for twenty minutes before anyone found out.

A scenario is checked against a schema, so a wrong step is wrong where it is written #

A scenario file used to be judged by the thing that replayed it, which means on a booked device, minutes into a run, after an emulator had been warmed and two apps signed in. Four kinds of mistake cost a whole run each, and one of them did not even fail: {"action":"wait","ms":6000} imported with no value at all and replayed as the 500ms default, so the pause a human wrote was silently a twelfth of what they asked for and the run went green.

  • schema/scenario.schema.json is JSON Schema draft 2020-12, and it travels inside the package so an editor can point at it without cloning anything.
  • It is GENERATED, not written: dart run tool/print_scenario_schema.dart reads [StepVocabulary] and emits it. A test fails when the two disagree, which is the only thing that keeps a schema honest as actions are added.
  • It refuses a key no importer reads, an action that needs a target and has none, a value on an action that reads no value, and a locator strategy that does not exist. Validated against every scenario in two real apps, 187 files, before release.

tryFind: a probe stops paying for an error message nobody reads #

[WidgetNotFoundError] builds its candidate list in its constructor, which is a full walk of the tree. A caller that catches the error and throws it away pays for all of it. Measured on a real Sites list: a find that hits costs 0.2s and a find that misses costs 12.1s. A client that walks indices until it runs out pays the miss on every walk, which is what made a list-ordering case slower than the code it was written to replace.

  • [SelfTestManager.tryFind] answers null instead of throwing. Anything that PROBES belongs there. Asserting still lists what was on screen, because there the message is the whole point.

The keys a step may carry are named beside the actions themselves #

The importer's rule about which keys a step may carry lived in the bridge while the actions lived in the vocabulary. A schema written against one and a refusal written against the other disagree the first time either moves.

  • [StepVocabulary.keysFor] holds it now, and the bridge reads it.
  • [StepKind.takesABudget] says which actions read a value as "wait this long for it": assertExists and assertAbsent. Without it, a schema derived from the vocabulary would have refused the budget the replayer supports and the importer's own error message recommends.
  • [SelfTestLocator]'s strategy names moved to a Flutter-free file so a generator, a linter or a schema tool can name a strategy without importing Flutter.

2.2.0 - 2026-09-22 #

Two features, both written because a real cross-surface case could not be expressed without them, and both proved against a live app before release.

A run can hand what it read to the phase that judges it #

The two halves of an end-to-end case used to share one process and pass a value through a static. That static is the only reason they had to be one process. A run now KEEPS what it reads, and writes it out, so the half that judges it can be another program in another language.

  • capture is a step: it names a widget with a locator and keeps the text it paints under a name of your choosing. It takes ONE look, like assertText, because the value field holds the fact's name and there is no budget written on the step to wait for.
  • An empty read is REFUSED rather than kept. An empty fact travels, matches everything downstream, and makes the assertion at the other end blame the wrong half for a value this step never read.
  • ${name} resolves in a step's value, in its locator and in its targetId. A name nobody captured is refused, and the refusal lists what WAS captured: sending the eleven characters ${siteName} into a field is a step that does exactly what it was told and reports PASSED.
  • Substitution never writes back into the stored script, so a script survives the run that used it.
  • dart run self_test:suite takes --facts-in and --facts-out. Facts are written after EVERY scenario, for the same reason the JUnit is: a run that dies half way still has to say what it learned.
  • Over the bridge: getFacts, setFacts, clearFacts. A non-string fact is refused rather than stringified.
  • Fixed before release: the replay loop settled on a step BEFORE resolving it, so a locator still reading ${siteName} named a widget that cannot exist and spent the whole settle budget every time. Measured with the fix reverted: 5.2s per step, on every run, and then the step works. Slow and correct is the hardest kind of waste to find. Resolution now happens ahead of the settle, in the normal and the repaired path alike, with a test that measures the clock.

A locator can be given a scope #

An index over the whole screen counts everything that happens to be built, so it moves when the data does. Measured on a real Sites list: the row titles sat at Text index 1, 7, 13, 18 and 24, because each row carries badges that are sometimes absent. The stride is that afternoon's data, not a layout, and nobody can write it down and have it keep working.

  • SelfTestLocator.within(scope), "within" on the wire: the search runs inside a resolved widget, so index 0 of Text inside ListTile 2 is the third row's title whatever the rows hold. Scopes nest.
  • This is what a geometric read ("the first text below the search box") approximates, without encoding the layout. A route transition puts both screens in the tree at once and the outgoing screen's rows sit below the incoming search box, which is exactly when geometry lies.
  • A scope that matches nothing resolves to nothing rather than quietly widening to the whole screen, and resolve names the missing scope instead of printing a list of the screen, which is the list that hides the answer.
  • Carried through at() and withValue(), so a ${fact} in a scoped locator stays scoped, and through JSON in both directions.
  • The report panel would have deleted it. The in-app editor rebuilds a locator from its form fields, so any edit would have dropped the scope in silence: the step would keep passing and start reading the whole screen. It now carries within through, clears it when a target is re-read off the screen, and shows it in the one-line rendering.

2.1.0 - 2026-09-22 #

Eight entries, every one of them found while driving two real apps against a live tenant, and every one of them carrying the test that pins it: the defects with a test that failed before the fix and passes after it, the guarantees with a test that would fail were the guarantee withdrawn.

A step whose value would be dropped is refused by name #

saveScript read a step's value from value, text or duration and dropped every other key without a word. A pause written {"action":"wait", "ms":6000} imported with no value at all and replayed as the 500ms default: the scenario asked for six seconds, got half of one, and still reported green. Measured, 20 Sep 2026: 2.34s wall clock against 8.51s for the same pause written the documented way. Eight scenarios in the suite this was found in had been written that way and none of their pauses had ever run. Importing such a step now fails and names the key.

An assertion can be told how long to wait #

assertExists and assertAbsent take an optional value in milliseconds and poll until the screen agrees or the budget runs out. Everything else keeps the five second settle replay already spends, so lengthening one assertion cannot slow the rest of a run. The cure for a step that fails on a busy device is not a longer pause: a pause long enough for the worst run is spent on every good one, and it is still a guess.

A scroll drags the list instead of sending a wheel #

scrollBy sent a single PointerScrollEvent, which is a mouse wheel, and a phone has none. It answered success for moving nothing, so a caller waiting for a row to come into view waited out its whole timeout with no error to read. Measured against a real device: three scrolls carrying a row's own locator each returned success and left the row at y 2675.0, to the pixel.

It now sends a drag, inside the scrollable the locator names or is, in strides no longer than the viewport, spending one deliberate move on the touch slop so the distance asked for is the distance the page travels, and holding still before lifting so the list does not fly past it. The sign is unchanged: a positive dy still means the page moves down, the way a wheel and a ScrollController both say it. A scroll may now name the list itself, not only a row inside it.

A drag is a step a scenario can write #

New action drag (alias dragAt), value x1,y1,x2,y2. The vocabulary had no action that carried a position: trigger presses the centre of whatever the locator finds, which on a slider is the middle of the track. A recording of a person dragging a slider to its minimum kept nothing at all, because a gesture that travelled was neither a tap nor a page movement and was dropped. The recorder now writes the path, and the replayer puts the finger back along it.

This is what a case about a derived value needs: an Amount(%) Standing slider with the standing tonnage read from it cannot be driven by a press.

A wait can watch what the app draws #

The bridge's wait takes a locator, not only a registered widget.

A wait may name the condition it already means #

saveScript refused any wait carrying a condition, because the key was not on the list of what that action reads. The bridge's own wait names what it is waiting for, so a step exported from a live session came back rejected by the import. A stored script can only let time pass, so condition: "duration" is now read as the harmless company it is. Any other condition is refused by name, with the assertion budget offered in its place, rather than being quietly turned into a pause that watches nothing.

A pause standing in front of an assertion is that assertion's budget #

Pinned by a test rather than left to be re-derived. A scenario written wait 2500 then exists X spends the whole 2.5 seconds on every run, including the runs where X was already drawn. Measured on two real suites on 21 Sep 2026: 229 pauses cost 531 seconds, 60.8% of all replay time, and 145 of them stood immediately in front of an assertion. The cure is not a shorter pause, which is the same guess made smaller: the pause is deleted and its milliseconds handed to the assertion, so a good run pays nothing and a slow one waits exactly as long as it used to. 2250 seconds across 137 scenarios were rewritten that way, which makes the import reading a budget out of an assertion load bearing for all of them.

A run record is claimed by identity, not by name #

self_test:suite attached the newest run stored under the scenario's name. The app appends its record a moment after the socket answers, so for the first instants of that lookup the newest run under that name is the PREVIOUS one: every report carried the duration, the app error count and the reportId the screenshots are pulled with of the run before it. It bites hardest on the resume pass after a failure, where a green verdict was showing the failed attempt's screenshots. The runner now reads which run ids exist before the scenario starts and waits for the one that was not there, falling back to its own stopwatch when the app is too old to keep records at all.

2.0.0 - 2026-09-16 #

Breaking #

Nothing was removed from the API. The major is here because four things a script or a CI job could already be reading now answer differently, and a minor version would have let them change under a caret.

  • goBack at the root stays in the app. It used to ask the platform to leave when there was nothing left to pop. On iOS that is a no-op, which is why it survived a year; on Android it sends the app to the launcher and the bridge dies with it, because the bridge lives inside the app. A run answered {"success": true, "popped": false}, the port stayed in LISTEN and the next command timed out, so every cheap check said the app was healthy. Every scenario here opens with five or six goBack calls to reach a known screen, so this fired on the first step of the first scenario and blamed whichever one happened to be next. It now answers popped: false and stays put. A script that used a trailing goBack to close the app has to close it some other way.

  • A failure is quoted once. The message that reaches the report page, the exported HTML and the JUnit <failure message> no longer arrives wrapped in quotes with its own quotes escaped. Anything matching on that string matches a different string now.

  • A step number counts from one. A failure used to report the app's index, so a scenario that broke on its fifth step read failed at step 4 of 5.

  • A step that names no widget reports no target. wait, goBack and mockChannel used to carry id: "" or the channel name in the field a widget id goes in.

Added #

  • A command line, so the suite runs in CI. dart run self_test:suite connects to a running app, installs each scenario in --dir as a script, runs it there and writes the verdict as JUnit XML. No model is in the loop and no tokens are spent: recording a scenario is the part a person does once, running it afterwards is the app executing its own steps. It runs on the Dart VM the Flutter install already provides, so a CI agent needs no Node, no Appium and no second toolchain. Exit code is 0 only when every scenario ran and passed, 1 when one failed or never ran, and 2 when the suite could not run at all, because a job that cannot tell "42 failed" from "42 never started" eventually learns to ignore both. --report-dir exports the app's own HTML report per scenario, screenshots inlined, plus an index. Jenkins, GitLab and GitHub recipes are in doc/ci/.

  • A scenario the suite never reached is written as skipped rather than left out. A sweep that died after 6 of 42 otherwise reports six passes and reads as a green build, which is how it happened twice in one afternoon.

  • Two things a pass cannot express travel in the XML: the count of steps the replayer had to repair, and the count of errors the app threw while the scenario passed anyway.

  • assertAbove, a step that fails unless its target is drawn higher up the screen than the text it names. It is the only way to write a rule about the order of a list: after a sort button is pressed the same rows are on screen whichever way the list came back, so assertExists on each of them passes even when the button does nothing at all. Spelled above for a hand-written script, and drawn in the check family on the report page.

  • A run record now carries what the app threw, not only what the replay refused. StepReport.appErrors holds the exceptions raised while that step ran, whatever its own verdict, and RunReport.appErrorCount summarises them so a list of runs can mark the green ones worth opening. The page shows them under the step that caused them and the export carries them with the stack, which is the whole of what a reader who cannot reproduce the run has.

    This was found on a real app: a draft with no scaffolds crashed during build and the scenario reported text: "Date and Time" is not there, which is true and says nothing about the crash. The worse case has no failure at all, and the run is green over an app that threw on every frame.

    FlutterError.onError and PlatformDispatcher.instance.onError are chained rather than replaced, so whatever the app already installed still runs, and both are put back when the run ends. Repeats of one error are merged and counted, since a broken build throws once per frame, and stacks are trimmed to twelve frames.

  • A mockChannel step can now carry a channels list. This lets one stored script answer platform-specific Pigeon channels without branching outside the app. The real case is image_picker: iOS asks image_picker_ios.ImagePickerApi.pickImage and expects one path, while Android asks image_picker_android.ImagePickerApi.pickImages and expects a list of paths. The replay installs both mocks and the platform consumes the one it calls. The SMART Handover tc273_handover_requires_a_photo scenario was rebuilt with that shape and passed on Android, 168/168.

Changed #

  • The panel and the controls inside the app are drawn in the same colours as the report page in the browser, and dark rather than white. This sheet is thrown over somebody else's product: a white panel over a white app left no line between what the app does and what the tool does, and the two halves of one tool disagreed about what a step looks like. SelfTestPalette holds the one set of colours and the report page builds its stylesheet from it, so a family added to StepVocabulary arrives in the browser and on the phone or in neither.

  • A step in the panel is read the way the page reads it: the action in the colour of its family, the locator in the words the locator uses, and the value on a surface of its own. It used to print trigger on and then the raw id, which for anything the recorder wrote is a line of JSON.

  • A list of scripts can be filtered by name. A device that has been driven for a week holds forty of them, and a flat list of forty is a list nobody reads.

  • A run says where it has got to while it is getting there: the script, the step and the count, the elapsed time and how many steps have been repaired, on a card over the app. A hundred and sixty eight steps take two minutes, and until now those two minutes had nothing on screen but a badge.

  • The example is a working host rather than a widget gallery. It starts the bridge outside release builds, keeps its scripts and its runs, and ships three scenarios, so flutter run followed by dart run self_test:suite shows a green suite, the report page and an exported report in about two minutes without an app of your own.

  • The README shows the tool rather than only describing it: the report page, a failed run beside the screen it failed on, the step editor resolving a name against the live app, and the panel drawn over the app. Every picture is of example/ in this repository, taken from a run of it, so what a reader sees is what flutter run in that directory gives them.

  • The setup instructions now say the host app needs uses-material-design: true. The panel is built from Material icons and an app that does not bundle that font draws every control in it as an empty box, which is how the example looked while it was being photographed. Declaring it in this package's pubspec instead was tried and changes nothing: the app's manifest is the one the tool reads.

Fixed #

  • The step a failure names is the one a human counts. The app answers with the index of the step it died on, and both the console line and the <failure message> in the JUnit XML printed that index as if it were the step number: a scenario that broke on its fifth step read failed at step 4 of 5. On a five-step login it is an off-by-one; on the 296-step scenario it was written against, it sends whoever opens the report to a different screen than the one that broke. The index is now converted where it arrives off the wire, and the field it lands in says so.

  • The example lets the bridge move between screens. It passed no navigator to SelfTestBridge and did not register the navigator's observer, so every goBack over the wire answered No BridgeNavigator is configured while the app carried on looking healthy. Its three scenarios each begin and end on the same screen, which is why nothing failed: a suite of more than one screen would have had its second scenario start wherever the first one left the app. The wiring is now the wiring the README's level 3 shows, and the example tests it.

  • A failed assertion is quoted once, not twice. AssertionError.toString() does not print its message: it prints Assertion failed: and then Error.safeToString(message), which wraps the message in quotes and escapes the ones inside it. A message that already began with Assertion failed: came out as Assertion failed: "Assertion failed: text: \"Passwordd\" ... is not there.", and that string is the whole of what a failure shows on the report page, in the exported HTML and in the JUnit <failure message> a CI server puts beside the red cross.

  • A step that names no widget no longer claims one. wait and goBack carry no locator and were recorded as id: "", and a mockChannel keeps the channel name where a widget id goes, so the record said the step had acted on a widget called dev.flutter.pigeon.image_picker_ios.ImagePickerApi.pickImage.

  • The controls no longer stand in the way of a replay. They stayed on screen throughout a run, at a fixed corner, hit-testing before the app under them: a step whose target happened to be beneath the button tapped the tool instead of the app and still reported itself done. Expanded, the cluster covered the whole screen with a barrier of its own and every remaining step landed on that. While a script is replaying the overlay now takes no pointers anywhere, and what it draws about the run ignores them.

1.0.0 #

First stable release. The public API below is the one this package intends to keep: a change that breaks it from here is a 2.0, not a patch. The 0.x line stopped at 0.1.0 on pub.flutter-io.cn, so every entry in this section is new to anyone installing the package rather than tracking the repository.

Breaking #

  • RecordingStore has one more method: deleteStep. A store that implements the interface has to add it, which is one line over whatever it already does for deleteScript. Without it there is no way to remove a step, so fixing one wrong step in a recording meant deleting the recording and capturing the whole session again.

  • Minimum Flutter is now 3.35.0 (Dart 3.9.0), raised from a declared 3.24.0 that was never true. lib/recording_fields passes widget properties that only exist from 3.35 (Switch.activeThumbColor and friends), so installing 0.1.0 on Flutter 3.24 produced eight compile errors inside this package. 3.35.0 is the oldest version the whole suite is verified against, and CI now runs against it on every push so the number stays honest.

  • The build_runner generator has moved to its own package, self_test_gen. Add it to dev_dependencies to keep using annotations. In exchange, the core package no longer depends on analyzer, source_gen, build or dart_style, so none of them reach your app any more. Its only dependencies are Flutter and meta.

  • Generated controller methods are camelCase, matching what the README always documented: tapLoginBtn() rather than tap_login_btn(). A private class such as _LoginFormState now generates LoginFormStateTestController, so it can be referenced from a test.

  • Do not write a part directive for the generated file and do not import it from its own source. The builder emits a standalone library; import it from your test. Importing it from its own source leaves that library unresolvable, and the generator then finds no annotations at all.

  • Recording persistence is an interface. RecordingStore and InMemoryRecordingStore replace DatabaseService, and hive_ce and path_provider are gone. Pass your own implementation to SelfTestManager().useRecordingStore() for recordings that survive a restart. initializeDatabase and clearDatabase are deprecated aliases for initializeRecordingStore and clearRecordings.

  • The recording model formerly called TestStep is now RecordedStep. TestStep remains the scenario command type it always was in the public API.

  • TestCodeGenerator returns strings instead of writing files, so the caller decides where generated code goes.

  • WidgetCatalog and FlowDiscovery are deprecated. Both only ever returned an empty map.

  • The bridge no longer depends on go_router. SelfTestBridge took a GoRouter? router, so a testing bridge that was meant to drive any Flutter app only navigated in apps that had picked one particular routing package. The parameter is now BridgeNavigator? navigator, an interface this package owns, with four members the bridge actually needs. Pass NavigatorStateBridgeNavigator(yourNavigatorKey) and it works in any app; a go_router app implements the interface in ten lines, and BridgeNavigator's doc comment gives that adapter in full.

    The parameter was renamed rather than kept as router, because it no longer takes a router. bridge.router is now bridge.navigator.

    Navigation commands that used to report {"success": true} while doing nothing, because no router was configured, now return an error saying so. getFlowGraph returns the routes the navigator reports rather than the empty map deprecated FlowDiscovery handed back.

  • The generator writes .self_test.g.dart, not .g.dart. Regenerate and update the import in your tests. source_gen:combining_builder claims .dart -> .g.dart, and every part-based generator (json_serializable, freezed, mobx, hive, drift) routes its output through it. Claiming the same file made build_runner refuse to start in the whole package, with Builders source_gen:combining_builder and self_test_gen:self_test outputs collide, taking the app's existing codegen down with it. Verified against a real app: a 75-file Flutter app using mobx_codegen and hive_ce_generator could not build at all with self_test_gen added, and builds with this change.

Added #

  • A run is coloured by what its steps do. Every action names a family, and the report page draws each family in its own colour: taps blue, typing teal, assertions violet, a channel mock orange, and getting around in grey so it recedes. A hundred and fifty rows of trigger differ by one word, and that word was the same colour as the rest of the line. The families come from StepVocabulary and reach the page over /report/api/vocabulary, so an action added to the list arrives already sorted, and a test fails if the page has no colour for a family the app names or a colour for one it does not. The verdict is on the edge of a row as well as in its tag, and an exported report is drawn the same way from the same list.

  • A script can be edited from the report page. Add, change, reorder and delete the steps of a stored script in the browser. SelfTestManager grew insertStep, editStep, deleteStepAt and moveStep; the page reaches them over /report/api/scripts/<id>/steps. A recording is a first draft, and until now one wrong step in a fifteen-step session meant recording the session again.

  • One step vocabulary, shared. StepVocabulary is the single list of what a step can be: the action, its other spellings, whether it needs a target, what its value means, and an example. The importer reads it, the editor is served it over /report/api/vocabulary, and a test compares it with the case labels of the replayer's own switch, so an action cannot be added to one and missed by the others. The README's step table is generated from it and a test fails until the file matches.

  • The editor is offered what is on screen. SelfTestManager.suggestTargets returns every widget the app can see with the strongest locator that finds it, a key before the words it paints, indexes filled in. /report/api/screen serves it and /report/api/resolve answers what a locator finds before a run depends on it: how many matches, which one the index picks, and whether that one has anything to press.

  • Import and export a script. GET /report/api/scripts/<id>/export answers one script as a file that carries no ids from the app that held it; POST /report/api/scripts/import takes one back, checked by the same rules as a suite installed over the socket. A file can also be dropped anywhere on the page.

  • The report page has been rebuilt: one icon set drawn on one grid, a step editor, and a test that reads the page as it is served and fails when the script reaches for an element the markup does not have. They fail together otherwise, silently, and what arrives in the browser is a header over three empty columns.

  • A run keeps a record of itself. Every replayed step is recorded with its outcome, its duration, the locator as recorded and the locator as repaired, and optionally a screenshot; RunReport and StepReport hold it and a RunReportStore keeps it on disk past the run. Before this a run left a verdict and nothing else: lastRunStatus was stamped on the script and overwrote the previous run, so a run that repaired two steps and passed was indistinguishable on screen from one that passed clean. The result knew the difference and nobody could see it.

  • Screenshots are a choice, not a policy. CaptureMode.failuresOnly is the default and photographs the first step, the last, and any failure; everyStep photographs all of them; never turns it off. Measured on a real app, a picture costs about 185ms and 47KB at pixelRatio 1.5, so photographing all 866 steps of a suite is a real cost to opt into rather than one to impose.

  • The report is served over the port the bridge already holds. GET /report answers an HTML page, /report/api/* answers JSON, and /report/shot/<run>/<step>.png answers a picture. The token is checked before the WebSocket upgrade is demanded, so these routes inherit the guard rather than adding a second one. It is served by the app because that is where the pictures are: on a simulator they can be copied out with simctl, on a device they cannot be reached at all.

  • The browser can drive the app. Run, stop, rename and delete a script, and delete a run, from the page. A run answers 202 immediately with running: true, because a scenario takes minutes and a browser that waited would time out and read as a failure.

  • A run says where it is while it runs. /report/api/status reports the script, the step index, the total and the elapsed time, announced before the step is attempted rather than after it completes: a step can wait seconds for its target, and reported on completion the step a run hangs on is the one step that never appears. The elapsed time is computed by the reader, not frozen into the progress, so it keeps moving during exactly the wait being watched.

  • The report exports. GET /report/export/<run> answers one self-contained HTML file with the pictures inlined and no links back to the server, so a failing run can be read on a machine that never had the app.

  • assertAbsent, because an absence is its own assertion. Five of the regression cases this package is measured against have an absence as their whole expected result ("the Done button is hidden", "scaffold 0001 is not listed"). The obvious workaround, asserting the text that replaces the control, passes on a screen showing both. An absence assertion also skips the settle loop: it names a widget that must never appear, so waiting for it to appear and settle could only spend the whole budget confirming what was already true on the first poll.

  • A recording keeps what the finger did to the page. A scroll is recorded from ScrollNotification rather than from pointer arithmetic, because the finger does not know where the page stopped: a drag hands off to a fling that keeps going after the finger leaves, and a page that reaches its end stops while the finger continues. Only drags are recorded (a form that scrolls a field into view when the keyboard opens does it again on replay), only net movement above a 1px floor (dragging a list already at the top stretches it and lets it go), and the step names the Scrollable, never a row, because the step exists to move the rows.

  • A recording keeps the file a native picker answered with. The answer is recorded as a mockChannel step placed before the tap that caused it, since installing it after the tap is installing it after the real picker already opened. The file is copied, not referenced: a picker answers with a path into a directory the app empties. Only answers that name a file are kept, because that is the one class of answer a replay cannot obtain again, while connectivity, location and preferences are answered the same way on the next run; without that rule four taps recorded as twenty four steps with a GPS position and a timestamp frozen inside them.

  • dragAt, the nameless sibling of tapAt, and renameTestScript. The first exists because nothing else exercises the scroll recorder: a drag with a locator records its own step and the recorder discards it as a duplicate, so until there was a nameless drag nothing could prove the recorder worked.

  • A mocked platform channel now answers the app. Call SelfTestWidgetsFlutterBinding.ensureInitialized() where the app called WidgetsFlutterBinding.ensureInitialized(), and the binding's messenger answers mocked channels in place of the platform: the camera on a simulator, a native picker, a permission dialog. ChannelMocks holds the table and a log of what the app asked. Plugins take the default messenger from the binding the moment it exists, so this is the only place a mock can be put in their path; in an app that already has a binding the call installs nothing and ChannelMocks.instance.intercepting stays false, so a caller can say so. Pigeon channels (image_picker, path_provider, url_launcher...) are mocked with ChannelMockKind.message: their reply is a one-element list.

  • The bridge's mockChannel tells the truth. It refuses, naming the binding, when nothing intercepts the app's calls, instead of reporting a mock it had registered for the wrong direction. codec: "message" mocks a Pigeon channel; channelLog lists the calls the app made on mocked channels and what answered each. The MCP tool flutter_mock_channel exposes both.

  • One run at a time, and cancelRun(). A replay lives in the app, not in whoever pressed Run, so a driver that dies mid-run left the app pressing its buttons for minutes while a fresh driver's taps all reported success and nothing happened. runTestScript now refuses a second run with a result that says one is under way, isRunning says so, and cancelRun() stops the run at the next step and puts the app back in the mode it was found in.

  • Universal locators. SelfTestLocator finds a widget by the text it paints, its ValueKey, its tooltip, its semantics label, its type, or a SelfTestableWidget id, with .at(n) to pick between duplicates. Resolution walks the element tree, so an app needs no wrapper, no annotation and no generated code to be driven. Locators are JSON in both directions, which is how the bridge and the MCP server will carry them.

  • Real pointer events. tap, doubleTap, longPress, dragFrom and scrollBy dispatch through GestureBinding.handlePointerEvent, the same entry point the engine uses. Hit testing runs, so a widget behind a dialog is not reachable and a disabled button swallows the tap. Before this, a driven tap invoked the app's callback directly and therefore passed on buttons the user could not even reach.

  • Real text entry. typeInto goes through EditableTextState.updateEditingValue, the method the soft keyboard calls, so input formatters run, onChanged fires and a TextFormField validates what was actually typed. It also resolves a field from the label beside it, which is how a person describes it.

  • describeScreen() returns every actionable widget on screen with its type, text, tooltip, rect and enabled state, for an agent deciding what to do next.

  • exists and isVisible are separate questions. A list keeps items built after they scroll away, and tapping one of those would land on whatever is drawn at those coordinates now, so a driven gesture refuses instead.

  • A recording keeps how each widget was addressed, not just an id. RecordedStep.locator carries the locator, so a session recorded on an app with no wrappers replays. A step recorded before this, or against an id, still replays through the id path.

  • recordAssertion records an assertion against a locator, so a recorded session can check something. Recording only actions produces a generated test that drives the app and asserts nothing, which passes on a blank screen.

  • The code generator emits the locator API and a testWidgets body that pumps between steps, because a driven tap is a real pointer event and nothing it changes is visible until the next frame. It also pumps the app for you when given an appExpression.

  • useClock lets a widget test hand the driver tester.pump, which is what a long press needs to be held rather than silently degrading to a tap.

  • The bridge speaks locators. Every command that named a widget by its registered id now also takes a locator object, and prefers it when both are sent:

    {"by": "text|key|id|semanticsLabel|type|tooltip",
     "value": "Sign in", "exact": true, "index": 0}
    

    tap, doubleTap, longPress, type/enterText, submit, drag, scroll and clear route through the manager's locator API when given one, so they reach widgets the app never registered. describeScreen, find, exists, isVisible and readText are new and take a locator only; describeScreen answers with WidgetSnapshot JSON. submit is a new command.

    A malformed locator is answered with what is wrong with it - an unknown by, a non-string value, a negative index - rather than a cast error from three layers down.

  • Recording, replay and test code generation, merged in from the feature/visual_tests line: a recording control panel, a recording-field registry, and TestCodeGenerator.

  • A DevTools extension and a standalone inspector under packages/.

  • SelfTestManager.captureScreenshotBytes() for platform-neutral capture, and clearScreenshotKey() so a disposed boundary cannot leave a stale key.

  • MIT license, replacing the previous custom terms.

  • CI covering every package, including a job that runs the README quickstart.

Security #

  • self_test is inert in a release build. Every action and every query is gated: no taps, no typing, no screenshots, no recording, no reading the widget tree, and no register of what is on screen is even built. A device farm that means to drive a signed build calls SelfTestManager.enableInReleaseBuilds() deliberately. Nothing flips it by accident, and debugSimulateReleaseBuild exists so the guard is covered by tests rather than asserted in a comment.
  • The bridge binds to loopback, not to every network interface. Until now it was InternetAddress.anyIPv4, so anyone on the same wifi could drive a colleague's debug build: read the tree, tap, type, photograph the screen. Pass host: InternetAddress.anyIPv4 to reach it from a real device, and it says out loud what that means.
  • The bridge requires a token. One is generated per instance and printed at startup, or the app supplies its own. A connection without it, or with a wrong one, gets an HTTP 403 before the WebSocket upgrade. The comparison is constant time.
  • The bridge refuses to start in a release build unless allowInReleaseBuilds: true is passed, because a bridge in a shipped app is a remote control for it.

Fixed #

  • A mock step exported with a script could not be imported again. The importer built a mock only from spread-out params, and a step read back from a script carries all three values in the single field a step has, so export was a one-way door for any script that answered a channel.
  • A step written by hand could name no widget at all. enterText against the empty locator resolved to nothing and reported a pass.

Every screenshot after the first was the same picture. With no boundary registered, capture walks the tree and photographs the first RenderRepaintBoundary it finds, and on a real app that one belonged to a route that had stopped repainting: a nine step run produced seven byte identical files, including the one labelled as the failure. The root the app is already wrapped in now installs the boundary, and capture waits while debugNeedsPaint, because toImage answers the last raster that was painted. The package's own docstring had warned about this.

  • A step that names no widget no longer waits for one. Before each step the replay waits for that step's target to appear and settle, with a five second budget. wait, goBack and mockChannel name no widget: the first two carry an empty target that resolves to nothing on every poll, and a mock carries a channel name, which looks like an ordinary target and will never resolve. Both spent the whole budget before running a step that was never about a widget. Nothing failed; the only symptom was the clock. Measured at 327 of the 866 steps in one suite, or 27 minutes of waiting for a widget no step mentions; one scenario went from 4.97s to 0.78s per step.

  • A field is named by the label beside it, not by what it shows. The recorder derived a name from the only text a widget painted, which in an empty field is its placeholder, and the replayer refuses a placeholder on purpose because it stops existing the moment the field is filled. So the recorder wrote a locator the replayer was designed to reject. Worse, the fallback was an index, and an index counted with a bottom sheet open is read again with it closed: a recorded EditableText index 1 typed into the date field and the run reported PASSED. A field is now named by the nearest word beside it, accepted only when asking the text input driver what that word means answers this field, which also settles which of two identical labels was meant.

  • Typing resolves to a field that can take a keystroke. A locator means the first match, and an app often draws a read only picture of a field over the real one: the first match was the picture, the guard refused it, and the run stopped with the field it wanted one match further along on the same screen. A typing step now skips matches with no field or a read only field, and when none of them accepts text the guard speaks as before.

  • A tap a guard refuses is retried once with the target scrolled into view. A recording has no scroll in it unless the finger made one, so a row recorded at the top of a form replays under a footer that covers it. The retry happens only after a refusal, never before a tap that was already going to land, and only once: when there is nothing to scroll, or scrolling does not free the target, the refusal stands.

  • A recorded picked file survives the next install. The step stored an absolute path into the app's data container, and iOS renames that container on every install while migrating the files: the file always survives, the path always dies. Three container names in two days. The app then received a path to nothing, which only failed much later inside the image decoder, far from the cause, and looked exactly like recording nothing at all. A step now stores self_test:kept/<name> and the replay joins the name to this install's directory, so a recording made two installs ago still runs.

  • A picker that names the same file twice is copied once. file_picker answers with a list of maps naming the file in both path and identifier as a file:// URL. The copy was memoized on the string, and two spellings are two strings, so the step ended up pointing at two different files. Reading the answer also walked only the strings of a list, so a map inside the list yielded no string at all, the whole call was discarded as noise, and the recording kept the tap that opened the gallery and nothing else.

  • A replay no longer writes into an open recording. The recorder watches pointers at the root and a replayed tap is a pointer like any other, so running a script while recording appended the replay's own taps to the recording; when the script being replayed was the one being recorded it doubled, and the run still reported PASSED. The only clue was the step count in the panel, by which time it is too late.

  • A scroll step is not repaired as though it were a tap. Recording the scroll gave the step a point, and self healing reads a point as a tap: a scroll was rewritten to be named after the row it exists to move away.

  • A picture waits for the app instead of a fixed delay. Capture happened 100ms after the step, which is a route transition and not a screen load, so two steps of a passing sixteen step run were photographs of a spinner on an otherwise blank screen. Measured on a real app, the same tap answered in under 170ms warm and had not finished in 500ms cold, so any hand picked number is short on one run and wasted on the other. Capture now waits for the next step's target to settle and for no progress indicator to be on screen, photographs anyway when the budget runs out and says so (capturedWhileBusy), and never waits at all for a failure (the screen that blocked it is the one worth photographing) or for a step nobody photographs. RefreshIndicator is explicitly not a progress indicator: it is the pull to refresh wrapper and sits in the tree of every list that can be pulled, loading or not, so counting it made a fully drawn list claim to be busy and cost the whole budget on every picture.

  • find answers the index it was given. It returned the same widget for index 0 and index 7, and returned a widget for an index that does not exist.

  • A command sent without a parameter it needs now says which parameter. Sixty five call sites read params['x'] as String directly, so omitting one answered type 'Null' is not a subtype of type 'String' in type cast, which names neither the parameter nor the command and reads like a crash in the bridge rather than a mistake in the request. Found by driving a real app: wait with no condition answered exactly that. The locator parameter already had this treatment; the scalar ones now do too, and a wrong type reports what was actually sent.

  • The docs no longer assume you are watching a flutter run console. The bridge token is announced through debugPrint, which reaches nobody under flutter build plus simctl launch or adb shell am start, which is how CI drives an app. Three places told you to use "the token the app printed"; they now tell you to pass your own.

  • getByText and getByRole read the element tree, like every other discovery command. They read the registered-node map instead, so on an app that never adopted SelfTestableWidget, which is every app before it adopts this package, they answered [] with no error. Found by driving a real production app: describeScreen returned a Text reading "Scaffolder", a locator exists on that same string answered true, and getByText('Scaffolder') answered nothing. getByRole also inferred the role by looking for "checkbox" or "switch" inside the widget's id string, so a button named checkbox_help was a checkbox and a real Checkbox with an unhelpful id was not; roles now map to the Flutter types that answer to them. Its name: filter searches the button's descendants, because a Flutter button paints no text of its own.

  • self_test_bridge no longer caps web_socket_channel at 2.x. It only uses WebSocketChannel and WebSocketChannel.connect, both unchanged in 3.x, so the caret pin bought nothing and cost the consumer a major version: adding the bridge to a real app downgraded its web_socket_channel from 3.0.3 to 2.4.0 and dragged build_runner's shelf stack down with it. Now >=2.4.0 <4.0.0, with the bridge's 60 tests verified on 3.0.3.

  • The bridge answered "done" before it had done anything. tap, type and eleven other commands called SelfTestManager.trigger and enterText without awaiting them. Those methods became async when they started dispatching real pointer events, so an agent that tapped and then read the screen was told the tap was finished and shown the screen from before it. unawaited_futures is now enforced in that package so it cannot come back.

  • Time-travel capture-on-interaction and test-step recording were implemented and then never called, so both recorded nothing. Recorded comments were dropped between the command and the step.

  • The widget rebuild profiler reported every rebuild as "Frame sample" instead of the reason it had worked out.

  • Starting memory profiling twice abandoned the first timer, leaving two running and interleaving their samples.

  • A recorded tap fired the app's handler twice. The recording wrappers called the driving API to record the action, which invoked the wrapper's callback, and then called the child's callback as well. Recording now records, and the child's own handler runs once; SelfTestableWidget.onTap is a fallback for a child that carries none.

  • Driving a widget wrapped in SelfTestableWidget with no matching recording builder recursed until the stack ran out, because the fallback wrapper called trigger from inside the tap it was handling.

  • Under integration_test, driven gestures did nothing at all and said nothing: that binding drops pointer events that did not come from a WidgetTester. The driver now detects it and names the one line that fixes it, binding.shouldPropagateDevicePointerEvents = true.

  • Annotations on class members are found. LibraryReader.annotatedWith only visits top-level declarations, so annotating a handler, which is what the README asks for, previously generated nothing at all.

  • Recorded values are escaped when generating test code. A value containing a quote, a dollar sign or a newline produced source that would not compile, or that picked up an unintended interpolation.

  • Two assertText steps in one recording no longer declare the same local twice, which made the generated file fail to compile.

  • SelfTestRoot no longer draws its recording controls in a widget test. They are a live overlay, so pumpAndSettle on a wrapped app never returned. The new showControls flag makes the choice explicit.

  • ScreenshotBoundary installs a real RepaintBoundary and registers it with the manager. It was a StatelessWidget that returned its child, so capture photographed whichever boundary happened to come first in the tree, or nothing.

  • The generator could not run on current Flutter at all. self_test_gen pinned analyzer below 9, and analyzer 7 throws Missing implementation of visitDotShorthandPropertyAccess while serialising any library that reaches the Flutter framework, which on Dart 3.12 and later is every library. build_runner build therefore worked on the declared floor, Flutter 3.35, and crashed on Flutter 3.47. It now takes analyzer >=8.1.1 <15.0.0, source_gen ^4.2.4 and build >=3.0.2 <5.0.0, which resolves to analyzer 10 on Dart 3.9 and analyzer 14 on Dart 3.13. Generated output is byte-identical under both. The quickstart CI job now runs on both ends of the range rather than on one of them, which is why nothing caught this.

  • The core package no longer publishes the rest of the monorepo inside itself. A .pubignore keeps the bridge, the generator, the inspector and the MCP server's Node tree out of the self_test archive, and keeps extension/devtools/build in, because that is how DevTools finds the panel.

0.1.0 #

  • Direct callback invocation for automated testing
  • Code generation with build_runner and annotations
  • SelfTestableWidget wrapper for manual setup
  • TestScenario and TestStep for multi-step test flows
  • Text assertion support for input field validation
  • Screenshot capture during test execution
  • Memory-safe test node management
  • Runtime testing in debug/profile builds
  • Test mode support for unit/integration tests
  • Generated test controllers with full type safety

0.0.1 #

  • Initial experimental release
0
likes
160
points
504
downloads

Documentation

API reference

Publisher

unverified uploader

Weekly Downloads

Automated regression testing for Flutter through direct callback invocation. Run tests in live app environments without UI simulation.

Repository (GitHub)
View/report issues

Topics

#testing #flutter #regression #automation

License

MIT (license)

Dependencies

flutter, meta

More

Packages that depend on self_test