self_test 2.3.0
self_test: ^2.3.0 copied to clipboard
Automated regression testing for Flutter through direct callback invocation. Run tests in live app environments without UI simulation.
2.3.0 - 2026-09-22 #
Three changes, all of them written because something upstream of a device run could be wrong for twenty minutes before anyone found out.
A scenario is checked against a schema, so a wrong step is wrong where it is written #
A scenario file used to be judged by the thing that replayed it, which means on
a booked device, minutes into a run, after an emulator had been warmed and two
apps signed in. Four kinds of mistake cost a whole run each, and one of them
did not even fail: {"action":"wait","ms":6000} imported with no value at all
and replayed as the 500ms default, so the pause a human wrote was silently a
twelfth of what they asked for and the run went green.
schema/scenario.schema.jsonis JSON Schema draft 2020-12, and it travels inside the package so an editor can point at it without cloning anything.- It is GENERATED, not written:
dart run tool/print_scenario_schema.dartreads [StepVocabulary] and emits it. A test fails when the two disagree, which is the only thing that keeps a schema honest as actions are added. - It refuses a key no importer reads, an action that needs a target and has none, a value on an action that reads no value, and a locator strategy that does not exist. Validated against every scenario in two real apps, 187 files, before release.
tryFind: a probe stops paying for an error message nobody reads #
[WidgetNotFoundError] builds its candidate list in its constructor, which is a
full walk of the tree. A caller that catches the error and throws it away pays
for all of it. Measured on a real Sites list: a find that hits costs 0.2s and
a find that misses costs 12.1s. A client that walks indices until it runs out
pays the miss on every walk, which is what made a list-ordering case slower than
the code it was written to replace.
- [SelfTestManager.tryFind] answers null instead of throwing. Anything that PROBES belongs there. Asserting still lists what was on screen, because there the message is the whole point.
The keys a step may carry are named beside the actions themselves #
The importer's rule about which keys a step may carry lived in the bridge while the actions lived in the vocabulary. A schema written against one and a refusal written against the other disagree the first time either moves.
- [StepVocabulary.keysFor] holds it now, and the bridge reads it.
- [StepKind.takesABudget] says which actions read a value as "wait this long
for it":
assertExistsandassertAbsent. Without it, a schema derived from the vocabulary would have refused the budget the replayer supports and the importer's own error message recommends. - [SelfTestLocator]'s strategy names moved to a Flutter-free file so a generator, a linter or a schema tool can name a strategy without importing Flutter.
2.2.0 - 2026-09-22 #
Two features, both written because a real cross-surface case could not be expressed without them, and both proved against a live app before release.
A run can hand what it read to the phase that judges it #
The two halves of an end-to-end case used to share one process and pass a value through a static. That static is the only reason they had to be one process. A run now KEEPS what it reads, and writes it out, so the half that judges it can be another program in another language.
captureis a step: it names a widget with a locator and keeps the text it paints under a name of your choosing. It takes ONE look, likeassertText, because the value field holds the fact's name and there is no budget written on the step to wait for.- An empty read is REFUSED rather than kept. An empty fact travels, matches everything downstream, and makes the assertion at the other end blame the wrong half for a value this step never read.
${name}resolves in a step's value, in its locator and in its targetId. A name nobody captured is refused, and the refusal lists what WAS captured: sending the eleven characters${siteName}into a field is a step that does exactly what it was told and reports PASSED.- Substitution never writes back into the stored script, so a script survives the run that used it.
dart run self_test:suitetakes--facts-inand--facts-out. Facts are written after EVERY scenario, for the same reason the JUnit is: a run that dies half way still has to say what it learned.- Over the bridge:
getFacts,setFacts,clearFacts. A non-string fact is refused rather than stringified. - Fixed before release: the replay loop settled on a step BEFORE resolving
it, so a locator still reading
${siteName}named a widget that cannot exist and spent the whole settle budget every time. Measured with the fix reverted: 5.2s per step, on every run, and then the step works. Slow and correct is the hardest kind of waste to find. Resolution now happens ahead of the settle, in the normal and the repaired path alike, with a test that measures the clock.
A locator can be given a scope #
An index over the whole screen counts everything that happens to be built, so
it moves when the data does. Measured on a real Sites list: the row titles sat
at Text index 1, 7, 13, 18 and 24, because each row carries badges that are
sometimes absent. The stride is that afternoon's data, not a layout, and
nobody can write it down and have it keep working.
SelfTestLocator.within(scope),"within"on the wire: the search runs inside a resolved widget, so index 0 ofTextinsideListTile2 is the third row's title whatever the rows hold. Scopes nest.- This is what a geometric read ("the first text below the search box") approximates, without encoding the layout. A route transition puts both screens in the tree at once and the outgoing screen's rows sit below the incoming search box, which is exactly when geometry lies.
- A scope that matches nothing resolves to nothing rather than quietly
widening to the whole screen, and
resolvenames the missing scope instead of printing a list of the screen, which is the list that hides the answer. - Carried through
at()andwithValue(), so a${fact}in a scoped locator stays scoped, and through JSON in both directions. - The report panel would have deleted it. The in-app editor rebuilds a
locator from its form fields, so any edit would have dropped the scope in
silence: the step would keep passing and start reading the whole screen. It
now carries
withinthrough, clears it when a target is re-read off the screen, and shows it in the one-line rendering.
2.1.0 - 2026-09-22 #
Eight entries, every one of them found while driving two real apps against a live tenant, and every one of them carrying the test that pins it: the defects with a test that failed before the fix and passes after it, the guarantees with a test that would fail were the guarantee withdrawn.
A step whose value would be dropped is refused by name #
saveScript read a step's value from value, text or duration and
dropped every other key without a word. A pause written {"action":"wait", "ms":6000} imported with no value at all and replayed as the 500ms default:
the scenario asked for six seconds, got half of one, and still reported green.
Measured, 20 Sep 2026: 2.34s wall clock against 8.51s for the same pause
written the documented way. Eight scenarios in the suite this was found in had
been written that way and none of their pauses had ever run. Importing such a
step now fails and names the key.
An assertion can be told how long to wait #
assertExists and assertAbsent take an optional value in milliseconds and
poll until the screen agrees or the budget runs out. Everything else keeps the
five second settle replay already spends, so lengthening one assertion cannot
slow the rest of a run. The cure for a step that fails on a busy device is not
a longer pause: a pause long enough for the worst run is spent on every good
one, and it is still a guess.
A scroll drags the list instead of sending a wheel #
scrollBy sent a single PointerScrollEvent, which is a mouse wheel, and a
phone has none. It answered success for moving nothing, so a caller waiting
for a row to come into view waited out its whole timeout with no error to
read. Measured against a real device: three scrolls carrying a row's own
locator each returned success and left the row at y 2675.0, to the pixel.
It now sends a drag, inside the scrollable the locator names or is, in strides
no longer than the viewport, spending one deliberate move on the touch slop so
the distance asked for is the distance the page travels, and holding still
before lifting so the list does not fly past it. The sign is unchanged: a
positive dy still means the page moves down, the way a wheel and a
ScrollController both say it. A scroll may now name the list itself, not
only a row inside it.
A drag is a step a scenario can write #
New action drag (alias dragAt), value x1,y1,x2,y2. The vocabulary had no
action that carried a position: trigger presses the centre of whatever the
locator finds, which on a slider is the middle of the track. A recording of a
person dragging a slider to its minimum kept nothing at all, because a gesture
that travelled was neither a tap nor a page movement and was dropped. The
recorder now writes the path, and the replayer puts the finger back along it.
This is what a case about a derived value needs: an Amount(%) Standing
slider with the standing tonnage read from it cannot be driven by a press.
A wait can watch what the app draws #
The bridge's wait takes a locator, not only a registered widget.
A wait may name the condition it already means #
saveScript refused any wait carrying a condition, because the key was not
on the list of what that action reads. The bridge's own wait names what it is
waiting for, so a step exported from a live session came back rejected by the
import. A stored script can only let time pass, so condition: "duration" is
now read as the harmless company it is. Any other condition is refused by name,
with the assertion budget offered in its place, rather than being quietly
turned into a pause that watches nothing.
A pause standing in front of an assertion is that assertion's budget #
Pinned by a test rather than left to be re-derived. A scenario written wait 2500 then exists X spends the whole 2.5 seconds on every run, including the
runs where X was already drawn. Measured on two real suites on 21 Sep 2026: 229
pauses cost 531 seconds, 60.8% of all replay time, and 145 of them stood
immediately in front of an assertion. The cure is not a shorter pause, which is
the same guess made smaller: the pause is deleted and its milliseconds handed
to the assertion, so a good run pays nothing and a slow one waits exactly as
long as it used to. 2250 seconds across 137 scenarios were rewritten that way,
which makes the import reading a budget out of an assertion load bearing for
all of them.
A run record is claimed by identity, not by name #
self_test:suite attached the newest run stored under the scenario's name. The
app appends its record a moment after the socket answers, so for the first
instants of that lookup the newest run under that name is the PREVIOUS one:
every report carried the duration, the app error count and the reportId the
screenshots are pulled with of the run before it. It bites hardest on the
resume pass after a failure, where a green verdict was showing the failed
attempt's screenshots. The runner now reads which run ids exist before the
scenario starts and waits for the one that was not there, falling back to its
own stopwatch when the app is too old to keep records at all.
2.0.0 - 2026-09-16 #
Breaking #
Nothing was removed from the API. The major is here because four things a script or a CI job could already be reading now answer differently, and a minor version would have let them change under a caret.
-
goBackat the root stays in the app. It used to ask the platform to leave when there was nothing left to pop. On iOS that is a no-op, which is why it survived a year; on Android it sends the app to the launcher and the bridge dies with it, because the bridge lives inside the app. A run answered{"success": true, "popped": false}, the port stayed inLISTENand the next command timed out, so every cheap check said the app was healthy. Every scenario here opens with five or sixgoBackcalls to reach a known screen, so this fired on the first step of the first scenario and blamed whichever one happened to be next. It now answerspopped: falseand stays put. A script that used a trailinggoBackto close the app has to close it some other way. -
A failure is quoted once. The message that reaches the report page, the exported HTML and the JUnit
<failure message>no longer arrives wrapped in quotes with its own quotes escaped. Anything matching on that string matches a different string now. -
A step number counts from one. A failure used to report the app's index, so a scenario that broke on its fifth step read
failed at step 4 of 5. -
A step that names no widget reports no target.
wait,goBackandmockChannelused to carryid: ""or the channel name in the field a widget id goes in.
Added #
-
A command line, so the suite runs in CI.
dart run self_test:suiteconnects to a running app, installs each scenario in--diras a script, runs it there and writes the verdict as JUnit XML. No model is in the loop and no tokens are spent: recording a scenario is the part a person does once, running it afterwards is the app executing its own steps. It runs on the Dart VM the Flutter install already provides, so a CI agent needs no Node, no Appium and no second toolchain. Exit code is 0 only when every scenario ran and passed, 1 when one failed or never ran, and 2 when the suite could not run at all, because a job that cannot tell "42 failed" from "42 never started" eventually learns to ignore both.--report-direxports the app's own HTML report per scenario, screenshots inlined, plus an index. Jenkins, GitLab and GitHub recipes are indoc/ci/. -
A scenario the suite never reached is written as
skippedrather than left out. A sweep that died after 6 of 42 otherwise reports six passes and reads as a green build, which is how it happened twice in one afternoon. -
Two things a pass cannot express travel in the XML: the count of steps the replayer had to repair, and the count of errors the app threw while the scenario passed anyway.
-
assertAbove, a step that fails unless its target is drawn higher up the screen than the text it names. It is the only way to write a rule about the order of a list: after a sort button is pressed the same rows are on screen whichever way the list came back, soassertExistson each of them passes even when the button does nothing at all. Spelledabovefor a hand-written script, and drawn in thecheckfamily on the report page. -
A run record now carries what the app threw, not only what the replay refused.
StepReport.appErrorsholds the exceptions raised while that step ran, whatever its own verdict, andRunReport.appErrorCountsummarises them so a list of runs can mark the green ones worth opening. The page shows them under the step that caused them and the export carries them with the stack, which is the whole of what a reader who cannot reproduce the run has.This was found on a real app: a draft with no scaffolds crashed during build and the scenario reported
text: "Date and Time" is not there, which is true and says nothing about the crash. The worse case has no failure at all, and the run is green over an app that threw on every frame.FlutterError.onErrorandPlatformDispatcher.instance.onErrorare chained rather than replaced, so whatever the app already installed still runs, and both are put back when the run ends. Repeats of one error are merged and counted, since a broken build throws once per frame, and stacks are trimmed to twelve frames. -
A
mockChannelstep can now carry achannelslist. This lets one stored script answer platform-specific Pigeon channels without branching outside the app. The real case isimage_picker: iOS asksimage_picker_ios.ImagePickerApi.pickImageand expects one path, while Android asksimage_picker_android.ImagePickerApi.pickImagesand expects a list of paths. The replay installs both mocks and the platform consumes the one it calls. The SMART Handovertc273_handover_requires_a_photoscenario was rebuilt with that shape and passed on Android, 168/168.
Changed #
-
The panel and the controls inside the app are drawn in the same colours as the report page in the browser, and dark rather than white. This sheet is thrown over somebody else's product: a white panel over a white app left no line between what the app does and what the tool does, and the two halves of one tool disagreed about what a step looks like.
SelfTestPaletteholds the one set of colours and the report page builds its stylesheet from it, so a family added toStepVocabularyarrives in the browser and on the phone or in neither. -
A step in the panel is read the way the page reads it: the action in the colour of its family, the locator in the words the locator uses, and the value on a surface of its own. It used to print
trigger onand then the raw id, which for anything the recorder wrote is a line of JSON. -
A list of scripts can be filtered by name. A device that has been driven for a week holds forty of them, and a flat list of forty is a list nobody reads.
-
A run says where it has got to while it is getting there: the script, the step and the count, the elapsed time and how many steps have been repaired, on a card over the app. A hundred and sixty eight steps take two minutes, and until now those two minutes had nothing on screen but a badge.
-
The example is a working host rather than a widget gallery. It starts the bridge outside release builds, keeps its scripts and its runs, and ships three scenarios, so
flutter runfollowed bydart run self_test:suiteshows a green suite, the report page and an exported report in about two minutes without an app of your own. -
The README shows the tool rather than only describing it: the report page, a failed run beside the screen it failed on, the step editor resolving a name against the live app, and the panel drawn over the app. Every picture is of
example/in this repository, taken from a run of it, so what a reader sees is whatflutter runin that directory gives them. -
The setup instructions now say the host app needs
uses-material-design: true. The panel is built from Material icons and an app that does not bundle that font draws every control in it as an empty box, which is how the example looked while it was being photographed. Declaring it in this package's pubspec instead was tried and changes nothing: the app's manifest is the one the tool reads.
Fixed #
-
The step a failure names is the one a human counts. The app answers with the index of the step it died on, and both the console line and the
<failure message>in the JUnit XML printed that index as if it were the step number: a scenario that broke on its fifth step readfailed at step 4 of 5. On a five-step login it is an off-by-one; on the 296-step scenario it was written against, it sends whoever opens the report to a different screen than the one that broke. The index is now converted where it arrives off the wire, and the field it lands in says so. -
The example lets the bridge move between screens. It passed no
navigatortoSelfTestBridgeand did not register the navigator's observer, so everygoBackover the wire answeredNo BridgeNavigator is configuredwhile the app carried on looking healthy. Its three scenarios each begin and end on the same screen, which is why nothing failed: a suite of more than one screen would have had its second scenario start wherever the first one left the app. The wiring is now the wiring the README's level 3 shows, and the example tests it. -
A failed assertion is quoted once, not twice.
AssertionError.toString()does not print its message: it printsAssertion failed:and thenError.safeToString(message), which wraps the message in quotes and escapes the ones inside it. A message that already began withAssertion failed:came out asAssertion failed: "Assertion failed: text: \"Passwordd\" ... is not there.", and that string is the whole of what a failure shows on the report page, in the exported HTML and in the JUnit<failure message>a CI server puts beside the red cross. -
A step that names no widget no longer claims one.
waitandgoBackcarry no locator and were recorded asid: "", and amockChannelkeeps the channel name where a widget id goes, so the record said the step had acted on a widget calleddev.flutter.pigeon.image_picker_ios.ImagePickerApi.pickImage. -
The controls no longer stand in the way of a replay. They stayed on screen throughout a run, at a fixed corner, hit-testing before the app under them: a step whose target happened to be beneath the button tapped the tool instead of the app and still reported itself done. Expanded, the cluster covered the whole screen with a barrier of its own and every remaining step landed on that. While a script is replaying the overlay now takes no pointers anywhere, and what it draws about the run ignores them.
1.0.0 #
First stable release. The public API below is the one this package intends to keep: a change that breaks it from here is a 2.0, not a patch. The 0.x line stopped at 0.1.0 on pub.flutter-io.cn, so every entry in this section is new to anyone installing the package rather than tracking the repository.
Breaking #
-
RecordingStorehas one more method:deleteStep. A store that implements the interface has to add it, which is one line over whatever it already does fordeleteScript. Without it there is no way to remove a step, so fixing one wrong step in a recording meant deleting the recording and capturing the whole session again. -
Minimum Flutter is now 3.35.0 (Dart 3.9.0), raised from a declared 3.24.0 that was never true.
lib/recording_fieldspasses widget properties that only exist from 3.35 (Switch.activeThumbColorand friends), so installing 0.1.0 on Flutter 3.24 produced eight compile errors inside this package. 3.35.0 is the oldest version the whole suite is verified against, and CI now runs against it on every push so the number stays honest. -
The
build_runnergenerator has moved to its own package,self_test_gen. Add it todev_dependenciesto keep using annotations. In exchange, the core package no longer depends onanalyzer,source_gen,buildordart_style, so none of them reach your app any more. Its only dependencies are Flutter andmeta. -
Generated controller methods are camelCase, matching what the README always documented:
tapLoginBtn()rather thantap_login_btn(). A private class such as_LoginFormStatenow generatesLoginFormStateTestController, so it can be referenced from a test. -
Do not write a
partdirective for the generated file and do not import it from its own source. The builder emits a standalone library; import it from your test. Importing it from its own source leaves that library unresolvable, and the generator then finds no annotations at all. -
Recording persistence is an interface.
RecordingStoreandInMemoryRecordingStorereplaceDatabaseService, andhive_ceandpath_providerare gone. Pass your own implementation toSelfTestManager().useRecordingStore()for recordings that survive a restart.initializeDatabaseandclearDatabaseare deprecated aliases forinitializeRecordingStoreandclearRecordings. -
The recording model formerly called
TestStepis nowRecordedStep.TestStepremains the scenario command type it always was in the public API. -
TestCodeGeneratorreturns strings instead of writing files, so the caller decides where generated code goes. -
WidgetCatalogandFlowDiscoveryare deprecated. Both only ever returned an empty map. -
The bridge no longer depends on go_router.
SelfTestBridgetook aGoRouter? router, so a testing bridge that was meant to drive any Flutter app only navigated in apps that had picked one particular routing package. The parameter is nowBridgeNavigator? navigator, an interface this package owns, with four members the bridge actually needs. PassNavigatorStateBridgeNavigator(yourNavigatorKey)and it works in any app; a go_router app implements the interface in ten lines, andBridgeNavigator's doc comment gives that adapter in full.The parameter was renamed rather than kept as
router, because it no longer takes a router.bridge.routeris nowbridge.navigator.Navigation commands that used to report
{"success": true}while doing nothing, because no router was configured, now return an error saying so.getFlowGraphreturns the routes the navigator reports rather than the empty map deprecatedFlowDiscoveryhanded back. -
The generator writes
.self_test.g.dart, not.g.dart. Regenerate and update the import in your tests.source_gen:combining_builderclaims.dart->.g.dart, and everypart-based generator (json_serializable, freezed, mobx, hive, drift) routes its output through it. Claiming the same file madebuild_runnerrefuse to start in the whole package, withBuilders source_gen:combining_builder and self_test_gen:self_test outputs collide, taking the app's existing codegen down with it. Verified against a real app: a 75-file Flutter app using mobx_codegen and hive_ce_generator could not build at all with self_test_gen added, and builds with this change.
Added #
-
A run is coloured by what its steps do. Every action names a family, and the report page draws each family in its own colour: taps blue, typing teal, assertions violet, a channel mock orange, and getting around in grey so it recedes. A hundred and fifty rows of
triggerdiffer by one word, and that word was the same colour as the rest of the line. The families come fromStepVocabularyand reach the page over/report/api/vocabulary, so an action added to the list arrives already sorted, and a test fails if the page has no colour for a family the app names or a colour for one it does not. The verdict is on the edge of a row as well as in its tag, and an exported report is drawn the same way from the same list. -
A script can be edited from the report page. Add, change, reorder and delete the steps of a stored script in the browser.
SelfTestManagergrewinsertStep,editStep,deleteStepAtandmoveStep; the page reaches them over/report/api/scripts/<id>/steps. A recording is a first draft, and until now one wrong step in a fifteen-step session meant recording the session again. -
One step vocabulary, shared.
StepVocabularyis the single list of what a step can be: the action, its other spellings, whether it needs a target, what its value means, and an example. The importer reads it, the editor is served it over/report/api/vocabulary, and a test compares it with thecaselabels of the replayer's own switch, so an action cannot be added to one and missed by the others. The README's step table is generated from it and a test fails until the file matches. -
The editor is offered what is on screen.
SelfTestManager.suggestTargetsreturns every widget the app can see with the strongest locator that finds it, a key before the words it paints, indexes filled in./report/api/screenserves it and/report/api/resolveanswers what a locator finds before a run depends on it: how many matches, which one the index picks, and whether that one has anything to press. -
Import and export a script.
GET /report/api/scripts/<id>/exportanswers one script as a file that carries no ids from the app that held it;POST /report/api/scripts/importtakes one back, checked by the same rules as a suite installed over the socket. A file can also be dropped anywhere on the page. -
The report page has been rebuilt: one icon set drawn on one grid, a step editor, and a test that reads the page as it is served and fails when the script reaches for an element the markup does not have. They fail together otherwise, silently, and what arrives in the browser is a header over three empty columns.
-
A run keeps a record of itself. Every replayed step is recorded with its outcome, its duration, the locator as recorded and the locator as repaired, and optionally a screenshot;
RunReportandStepReporthold it and aRunReportStorekeeps it on disk past the run. Before this a run left a verdict and nothing else:lastRunStatuswas stamped on the script and overwrote the previous run, so a run that repaired two steps and passed was indistinguishable on screen from one that passed clean. The result knew the difference and nobody could see it. -
Screenshots are a choice, not a policy.
CaptureMode.failuresOnlyis the default and photographs the first step, the last, and any failure;everyStepphotographs all of them;neverturns it off. Measured on a real app, a picture costs about 185ms and 47KB atpixelRatio1.5, so photographing all 866 steps of a suite is a real cost to opt into rather than one to impose. -
The report is served over the port the bridge already holds.
GET /reportanswers an HTML page,/report/api/*answers JSON, and/report/shot/<run>/<step>.pnganswers a picture. The token is checked before the WebSocket upgrade is demanded, so these routes inherit the guard rather than adding a second one. It is served by the app because that is where the pictures are: on a simulator they can be copied out withsimctl, on a device they cannot be reached at all. -
The browser can drive the app. Run, stop, rename and delete a script, and delete a run, from the page. A run answers
202immediately withrunning: true, because a scenario takes minutes and a browser that waited would time out and read as a failure. -
A run says where it is while it runs.
/report/api/statusreports the script, the step index, the total and the elapsed time, announced before the step is attempted rather than after it completes: a step can wait seconds for its target, and reported on completion the step a run hangs on is the one step that never appears. The elapsed time is computed by the reader, not frozen into the progress, so it keeps moving during exactly the wait being watched. -
The report exports.
GET /report/export/<run>answers one self-contained HTML file with the pictures inlined and no links back to the server, so a failing run can be read on a machine that never had the app. -
assertAbsent, because an absence is its own assertion. Five of the regression cases this package is measured against have an absence as their whole expected result ("the Done button is hidden", "scaffold 0001 is not listed"). The obvious workaround, asserting the text that replaces the control, passes on a screen showing both. An absence assertion also skips the settle loop: it names a widget that must never appear, so waiting for it to appear and settle could only spend the whole budget confirming what was already true on the first poll. -
A recording keeps what the finger did to the page. A scroll is recorded from
ScrollNotificationrather than from pointer arithmetic, because the finger does not know where the page stopped: a drag hands off to a fling that keeps going after the finger leaves, and a page that reaches its end stops while the finger continues. Only drags are recorded (a form that scrolls a field into view when the keyboard opens does it again on replay), only net movement above a 1px floor (dragging a list already at the top stretches it and lets it go), and the step names theScrollable, never a row, because the step exists to move the rows. -
A recording keeps the file a native picker answered with. The answer is recorded as a
mockChannelstep placed before the tap that caused it, since installing it after the tap is installing it after the real picker already opened. The file is copied, not referenced: a picker answers with a path into a directory the app empties. Only answers that name a file are kept, because that is the one class of answer a replay cannot obtain again, while connectivity, location and preferences are answered the same way on the next run; without that rule four taps recorded as twenty four steps with a GPS position and a timestamp frozen inside them. -
dragAt, the nameless sibling oftapAt, andrenameTestScript. The first exists because nothing else exercises the scroll recorder: a drag with a locator records its own step and the recorder discards it as a duplicate, so until there was a nameless drag nothing could prove the recorder worked. -
A mocked platform channel now answers the app. Call
SelfTestWidgetsFlutterBinding.ensureInitialized()where the app calledWidgetsFlutterBinding.ensureInitialized(), and the binding's messenger answers mocked channels in place of the platform: the camera on a simulator, a native picker, a permission dialog.ChannelMocksholds the table and a log of what the app asked. Plugins take the default messenger from the binding the moment it exists, so this is the only place a mock can be put in their path; in an app that already has a binding the call installs nothing andChannelMocks.instance.interceptingstays false, so a caller can say so. Pigeon channels (image_picker,path_provider,url_launcher...) are mocked withChannelMockKind.message: their reply is a one-element list. -
The bridge's
mockChanneltells the truth. It refuses, naming the binding, when nothing intercepts the app's calls, instead of reporting a mock it had registered for the wrong direction.codec: "message"mocks a Pigeon channel;channelLoglists the calls the app made on mocked channels and what answered each. The MCP toolflutter_mock_channelexposes both. -
One run at a time, and
cancelRun(). A replay lives in the app, not in whoever pressed Run, so a driver that dies mid-run left the app pressing its buttons for minutes while a fresh driver's taps all reported success and nothing happened.runTestScriptnow refuses a second run with a result that says one is under way,isRunningsays so, andcancelRun()stops the run at the next step and puts the app back in the mode it was found in. -
Universal locators.
SelfTestLocatorfinds a widget by the text it paints, itsValueKey, its tooltip, its semantics label, its type, or aSelfTestableWidgetid, with.at(n)to pick between duplicates. Resolution walks the element tree, so an app needs no wrapper, no annotation and no generated code to be driven. Locators are JSON in both directions, which is how the bridge and the MCP server will carry them. -
Real pointer events.
tap,doubleTap,longPress,dragFromandscrollBydispatch throughGestureBinding.handlePointerEvent, the same entry point the engine uses. Hit testing runs, so a widget behind a dialog is not reachable and a disabled button swallows the tap. Before this, a driven tap invoked the app's callback directly and therefore passed on buttons the user could not even reach. -
Real text entry.
typeIntogoes throughEditableTextState.updateEditingValue, the method the soft keyboard calls, so input formatters run,onChangedfires and aTextFormFieldvalidates what was actually typed. It also resolves a field from the label beside it, which is how a person describes it. -
describeScreen()returns every actionable widget on screen with its type, text, tooltip, rect and enabled state, for an agent deciding what to do next. -
existsandisVisibleare separate questions. A list keeps items built after they scroll away, and tapping one of those would land on whatever is drawn at those coordinates now, so a driven gesture refuses instead. -
A recording keeps how each widget was addressed, not just an id.
RecordedStep.locatorcarries the locator, so a session recorded on an app with no wrappers replays. A step recorded before this, or against an id, still replays through the id path. -
recordAssertionrecords an assertion against a locator, so a recorded session can check something. Recording only actions produces a generated test that drives the app and asserts nothing, which passes on a blank screen. -
The code generator emits the locator API and a
testWidgetsbody that pumps between steps, because a driven tap is a real pointer event and nothing it changes is visible until the next frame. It also pumps the app for you when given anappExpression. -
useClocklets a widget test hand the drivertester.pump, which is what a long press needs to be held rather than silently degrading to a tap. -
The bridge speaks locators. Every command that named a widget by its registered id now also takes a
locatorobject, and prefers it when both are sent:{"by": "text|key|id|semanticsLabel|type|tooltip", "value": "Sign in", "exact": true, "index": 0}tap,doubleTap,longPress,type/enterText,submit,drag,scrollandclearroute through the manager's locator API when given one, so they reach widgets the app never registered.describeScreen,find,exists,isVisibleandreadTextare new and take a locator only;describeScreenanswers withWidgetSnapshotJSON.submitis a new command.A malformed locator is answered with what is wrong with it - an unknown
by, a non-stringvalue, a negativeindex- rather than a cast error from three layers down. -
Recording, replay and test code generation, merged in from the
feature/visual_testsline: a recording control panel, a recording-field registry, andTestCodeGenerator. -
A DevTools extension and a standalone inspector under
packages/. -
SelfTestManager.captureScreenshotBytes()for platform-neutral capture, andclearScreenshotKey()so a disposed boundary cannot leave a stale key. -
MIT license, replacing the previous custom terms.
-
CI covering every package, including a job that runs the README quickstart.
Security #
- self_test is inert in a release build. Every action and every query is
gated: no taps, no typing, no screenshots, no recording, no reading the
widget tree, and no register of what is on screen is even built. A device
farm that means to drive a signed build calls
SelfTestManager.enableInReleaseBuilds()deliberately. Nothing flips it by accident, anddebugSimulateReleaseBuildexists so the guard is covered by tests rather than asserted in a comment. - The bridge binds to loopback, not to every network interface. Until now
it was
InternetAddress.anyIPv4, so anyone on the same wifi could drive a colleague's debug build: read the tree, tap, type, photograph the screen. Passhost: InternetAddress.anyIPv4to reach it from a real device, and it says out loud what that means. - The bridge requires a token. One is generated per instance and printed at startup, or the app supplies its own. A connection without it, or with a wrong one, gets an HTTP 403 before the WebSocket upgrade. The comparison is constant time.
- The bridge refuses to start in a release build unless
allowInReleaseBuilds: trueis passed, because a bridge in a shipped app is a remote control for it.
Fixed #
- A mock step exported with a script could not be imported again. The importer built a mock only from spread-out params, and a step read back from a script carries all three values in the single field a step has, so export was a one-way door for any script that answered a channel.
- A step written by hand could name no widget at all.
enterTextagainst the empty locator resolved to nothing and reported a pass.
Every screenshot after the first was the same picture. With no boundary
registered, capture walks the tree and photographs the first
RenderRepaintBoundary it finds, and on a real app that one belonged to a
route that had stopped repainting: a nine step run produced seven
byte identical files, including the one labelled as the failure. The root the
app is already wrapped in now installs the boundary, and capture waits while
debugNeedsPaint, because toImage answers the last raster that was
painted. The package's own docstring had warned about this.
-
A step that names no widget no longer waits for one. Before each step the replay waits for that step's target to appear and settle, with a five second budget.
wait,goBackandmockChannelname no widget: the first two carry an empty target that resolves to nothing on every poll, and a mock carries a channel name, which looks like an ordinary target and will never resolve. Both spent the whole budget before running a step that was never about a widget. Nothing failed; the only symptom was the clock. Measured at 327 of the 866 steps in one suite, or 27 minutes of waiting for a widget no step mentions; one scenario went from 4.97s to 0.78s per step. -
A field is named by the label beside it, not by what it shows. The recorder derived a name from the only text a widget painted, which in an empty field is its placeholder, and the replayer refuses a placeholder on purpose because it stops existing the moment the field is filled. So the recorder wrote a locator the replayer was designed to reject. Worse, the fallback was an index, and an index counted with a bottom sheet open is read again with it closed: a recorded
EditableText index 1typed into the date field and the run reported PASSED. A field is now named by the nearest word beside it, accepted only when asking the text input driver what that word means answers this field, which also settles which of two identical labels was meant. -
Typing resolves to a field that can take a keystroke. A locator means the first match, and an app often draws a read only picture of a field over the real one: the first match was the picture, the guard refused it, and the run stopped with the field it wanted one match further along on the same screen. A typing step now skips matches with no field or a read only field, and when none of them accepts text the guard speaks as before.
-
A tap a guard refuses is retried once with the target scrolled into view. A recording has no scroll in it unless the finger made one, so a row recorded at the top of a form replays under a footer that covers it. The retry happens only after a refusal, never before a tap that was already going to land, and only once: when there is nothing to scroll, or scrolling does not free the target, the refusal stands.
-
A recorded picked file survives the next install. The step stored an absolute path into the app's data container, and iOS renames that container on every install while migrating the files: the file always survives, the path always dies. Three container names in two days. The app then received a path to nothing, which only failed much later inside the image decoder, far from the cause, and looked exactly like recording nothing at all. A step now stores
self_test:kept/<name>and the replay joins the name to this install's directory, so a recording made two installs ago still runs. -
A picker that names the same file twice is copied once.
file_pickeranswers with a list of maps naming the file in bothpathandidentifieras afile://URL. The copy was memoized on the string, and two spellings are two strings, so the step ended up pointing at two different files. Reading the answer also walked only the strings of a list, so a map inside the list yielded no string at all, the whole call was discarded as noise, and the recording kept the tap that opened the gallery and nothing else. -
A replay no longer writes into an open recording. The recorder watches pointers at the root and a replayed tap is a pointer like any other, so running a script while recording appended the replay's own taps to the recording; when the script being replayed was the one being recorded it doubled, and the run still reported PASSED. The only clue was the step count in the panel, by which time it is too late.
-
A scroll step is not repaired as though it were a tap. Recording the scroll gave the step a point, and self healing reads a point as a tap: a scroll was rewritten to be named after the row it exists to move away.
-
A picture waits for the app instead of a fixed delay. Capture happened 100ms after the step, which is a route transition and not a screen load, so two steps of a passing sixteen step run were photographs of a spinner on an otherwise blank screen. Measured on a real app, the same tap answered in under 170ms warm and had not finished in 500ms cold, so any hand picked number is short on one run and wasted on the other. Capture now waits for the next step's target to settle and for no progress indicator to be on screen, photographs anyway when the budget runs out and says so (
capturedWhileBusy), and never waits at all for a failure (the screen that blocked it is the one worth photographing) or for a step nobody photographs.RefreshIndicatoris explicitly not a progress indicator: it is the pull to refresh wrapper and sits in the tree of every list that can be pulled, loading or not, so counting it made a fully drawn list claim to be busy and cost the whole budget on every picture. -
findanswers the index it was given. It returned the same widget for index 0 and index 7, and returned a widget for an index that does not exist. -
A command sent without a parameter it needs now says which parameter. Sixty five call sites read
params['x'] as Stringdirectly, so omitting one answeredtype 'Null' is not a subtype of type 'String' in type cast, which names neither the parameter nor the command and reads like a crash in the bridge rather than a mistake in the request. Found by driving a real app:waitwith noconditionanswered exactly that. The locator parameter already had this treatment; the scalar ones now do too, and a wrong type reports what was actually sent. -
The docs no longer assume you are watching a
flutter runconsole. The bridge token is announced throughdebugPrint, which reaches nobody underflutter buildplussimctl launchoradb shell am start, which is how CI drives an app. Three places told you to use "the token the app printed"; they now tell you to pass your own. -
getByTextandgetByRoleread the element tree, like every other discovery command. They read the registered-node map instead, so on an app that never adoptedSelfTestableWidget, which is every app before it adopts this package, they answered[]with no error. Found by driving a real production app:describeScreenreturned aTextreading "Scaffolder", a locatorexistson that same string answered true, andgetByText('Scaffolder')answered nothing.getByRolealso inferred the role by looking for "checkbox" or "switch" inside the widget's id string, so a button namedcheckbox_helpwas a checkbox and a realCheckboxwith an unhelpful id was not; roles now map to the Flutter types that answer to them. Itsname:filter searches the button's descendants, because a Flutter button paints no text of its own. -
self_test_bridgeno longer capsweb_socket_channelat 2.x. It only usesWebSocketChannelandWebSocketChannel.connect, both unchanged in 3.x, so the caret pin bought nothing and cost the consumer a major version: adding the bridge to a real app downgraded itsweb_socket_channelfrom 3.0.3 to 2.4.0 and draggedbuild_runner's shelf stack down with it. Now>=2.4.0 <4.0.0, with the bridge's 60 tests verified on 3.0.3. -
The bridge answered "done" before it had done anything.
tap,typeand eleven other commands calledSelfTestManager.triggerandenterTextwithout awaiting them. Those methods became async when they started dispatching real pointer events, so an agent that tapped and then read the screen was told the tap was finished and shown the screen from before it.unawaited_futuresis now enforced in that package so it cannot come back. -
Time-travel capture-on-interaction and test-step recording were implemented and then never called, so both recorded nothing. Recorded comments were dropped between the command and the step.
-
The widget rebuild profiler reported every rebuild as "Frame sample" instead of the reason it had worked out.
-
Starting memory profiling twice abandoned the first timer, leaving two running and interleaving their samples.
-
A recorded tap fired the app's handler twice. The recording wrappers called the driving API to record the action, which invoked the wrapper's callback, and then called the child's callback as well. Recording now records, and the child's own handler runs once;
SelfTestableWidget.onTapis a fallback for a child that carries none. -
Driving a widget wrapped in
SelfTestableWidgetwith no matching recording builder recursed until the stack ran out, because the fallback wrapper calledtriggerfrom inside the tap it was handling. -
Under
integration_test, driven gestures did nothing at all and said nothing: that binding drops pointer events that did not come from aWidgetTester. The driver now detects it and names the one line that fixes it,binding.shouldPropagateDevicePointerEvents = true. -
Annotations on class members are found.
LibraryReader.annotatedWithonly visits top-level declarations, so annotating a handler, which is what the README asks for, previously generated nothing at all. -
Recorded values are escaped when generating test code. A value containing a quote, a dollar sign or a newline produced source that would not compile, or that picked up an unintended interpolation.
-
Two
assertTextsteps in one recording no longer declare the same local twice, which made the generated file fail to compile. -
SelfTestRootno longer draws its recording controls in a widget test. They are a live overlay, sopumpAndSettleon a wrapped app never returned. The newshowControlsflag makes the choice explicit. -
ScreenshotBoundaryinstalls a realRepaintBoundaryand registers it with the manager. It was aStatelessWidgetthat returned its child, so capture photographed whichever boundary happened to come first in the tree, or nothing. -
The generator could not run on current Flutter at all.
self_test_genpinnedanalyzerbelow 9, and analyzer 7 throwsMissing implementation of visitDotShorthandPropertyAccesswhile serialising any library that reaches the Flutter framework, which on Dart 3.12 and later is every library.build_runner buildtherefore worked on the declared floor, Flutter 3.35, and crashed on Flutter 3.47. It now takesanalyzer >=8.1.1 <15.0.0,source_gen ^4.2.4andbuild >=3.0.2 <5.0.0, which resolves to analyzer 10 on Dart 3.9 and analyzer 14 on Dart 3.13. Generated output is byte-identical under both. The quickstart CI job now runs on both ends of the range rather than on one of them, which is why nothing caught this. -
The core package no longer publishes the rest of the monorepo inside itself. A
.pubignorekeeps the bridge, the generator, the inspector and the MCP server's Node tree out of theself_testarchive, and keepsextension/devtools/buildin, because that is how DevTools finds the panel.
0.1.0 #
- Direct callback invocation for automated testing
- Code generation with
build_runnerand annotations SelfTestableWidgetwrapper for manual setupTestScenarioandTestStepfor multi-step test flows- Text assertion support for input field validation
- Screenshot capture during test execution
- Memory-safe test node management
- Runtime testing in debug/profile builds
- Test mode support for unit/integration tests
- Generated test controllers with full type safety
0.0.1 #
- Initial experimental release