Codifies the ticket-writing, human-validation, wave-dispatch, worktree-merge, emulator-verification, and APK-commit loop this repo's UI-redesign and feedback batches were built with, written in ASD-STE100 so ticket files and subagent prompts stay unambiguous for a fresh subagent with no conversation history.
271 lines
16 KiB
Markdown
271 lines
16 KiB
Markdown
---
|
||
name: autopilot-build
|
||
description: "Turn a markdown requirements file into implemented, tested, emulator-verified, committed Flutter changes for this repo (rippr-flutter) — breaks requirements into independently-completable tickets, pauses for human approval, then dispatches implementation to isolated subagents, merges, verifies on the Android emulator, and commits a fresh APK."
|
||
---
|
||
|
||
# Autopilot Build
|
||
|
||
Runs the same disciplined loop this repo's UI-redesign (`docs/ui-redesign/`) and
|
||
post-launch-feedback (`docs/feedback/`) batches were built with, end to end, from a
|
||
single requirements document. It is slow and thorough by design — every phase below
|
||
exists because skipping it caused a real problem in a prior run of this loop.
|
||
|
||
## Writing style: ASD-STE100 (Simplified Technical English)
|
||
|
||
Every ticket file, `Outcome` section, and subagent dispatch prompt this skill produces
|
||
is written in **ASD-STE100** — the controlled-language standard aerospace maintenance
|
||
manuals use specifically to remove ambiguity for a reader with no shared context. That
|
||
is exactly the constraint a fresh subagent with zero conversation history is under, so
|
||
the standard's core rules apply directly, not as a style nicety:
|
||
|
||
- **One instruction, one sentence.** Never join two instructions with "and" or a comma.
|
||
"Remove the panel. Disconnect the cable." not "Remove the panel and disconnect the
|
||
cable."
|
||
- **Active voice, imperative for instructions.** "Add the constant to `ride_map.dart`."
|
||
not "The constant should be added to `ride_map.dart`."
|
||
- **Keep sentences short** (roughly 20 words or fewer). Split a long sentence at its
|
||
natural joints rather than adding a subordinate clause.
|
||
- **Use one word for one meaning, every time.** If a ticket calls it "the overflow
|
||
menu" in one place, never call the same thing "the options menu" or "the kebab menu"
|
||
later in the same ticket. Inconsistent naming is the single most common way a
|
||
fresh-context reader misidentifies which code you mean.
|
||
- **Avoid long noun strings.** "the HUD widget grid collision check" is four nouns
|
||
stacked — prefer "the check that finds a collision in the HUD widget grid."
|
||
- **Avoid ambiguous pronouns.** Repeat the noun instead of "it"/"this"/"that" whenever
|
||
more than one candidate referent exists in the surrounding sentences — the exact
|
||
failure mode a fresh subagent can't recover from is guessing which prior noun "it"
|
||
meant.
|
||
- **Prefer concrete, specific verbs over vague ones.** "Reject the move" not "handle
|
||
the move"; "return the cached value" not "deal with the cached value."
|
||
- **State conditions before actions.** "If the route has fewer than two waypoints,
|
||
disable the toggle." not "Disable the toggle if the route has fewer than two
|
||
waypoints" — STE100 puts the condition first so the reader isn't committed to an
|
||
action before learning it might not apply.
|
||
|
||
This does not apply to direct quotations of existing code or user feedback (quote those
|
||
verbatim, exactly as written) — only to the ticket prose you write around them.
|
||
|
||
## Input
|
||
|
||
One argument: a path to a markdown requirements file (e.g. `docs/FEEDBACK.md`, or
|
||
whatever the user hands you). If no path is given, ask for one — don't guess at
|
||
requirements from conversation context alone; the file is the source of truth this
|
||
whole run is grounded in.
|
||
|
||
## Phase 0 — Sanity check before starting
|
||
|
||
- `flutter analyze` clean and `flutter test` green. Record the passing count — every
|
||
later phase's bar is "still clean, count only goes up from here."
|
||
- `git status` clean (or confirm with the user what to do with any dirty state — don't
|
||
silently work on top of someone else's uncommitted changes).
|
||
- Confirm a remote exists if this run is expected to end in something pushable — this
|
||
skill never pushes itself (see Phase 8), but say so up front if there's no remote.
|
||
|
||
## Phase 1 — Break requirements into tickets
|
||
|
||
Read the requirements file in full. Read enough of the current codebase (grep, targeted
|
||
file reads, and/or `Explore`-type subagents if the surface area is large or unfamiliar)
|
||
to ground each requirement in real files, not guesses. For anything with real design
|
||
ambiguity (a new data model, a rendering bug with no obvious cause, an architecture
|
||
change), launch a `Plan`-type subagent to work through the approach *before* you commit
|
||
to a ticket list — cheaper to redesign a paragraph now than a shipped ticket later.
|
||
|
||
Split the requirements into a small number of tickets, each one:
|
||
- **Independently completable** — a fresh subagent with zero conversation history could
|
||
read the ticket file alone and implement it correctly. This is the single most
|
||
important property. If a ticket needs context only available in this conversation,
|
||
the ticket is incomplete, not the subagent.
|
||
- **Scoped to avoid unnecessary file overlap** with its siblings where possible. Where
|
||
overlap is unavoidable (two tickets both need to touch the same screen), note the
|
||
dependency explicitly — see Phase 2.
|
||
|
||
**Do not write full ticket files yet.** Recite the list to the user first: one line
|
||
per ticket (title, one-sentence scope, rough size S/M/L). This mirrors the real
|
||
worked example: five tickets were named and sized in one message, and only written out
|
||
in full after the user confirmed the list. Skipping this step means discovering scope
|
||
disagreements after hours of writing detail nobody asked for.
|
||
|
||
## Phase 2 — Human validation (hard stop)
|
||
|
||
Do not proceed to Phase 3 without an explicit go-ahead on the ticket list. Use
|
||
`AskUserQuestion` if there's a real fork to resolve (e.g. "should FB-04 and FB-05 be one
|
||
ticket or two?"), otherwise just present the recited list and wait for the user's next
|
||
message. Adjust scope, splitting, or sizing based on their response before moving on —
|
||
this is the point where it's cheap to change direction.
|
||
|
||
## Phase 3 — Deep-dive each ticket and write full ticket files
|
||
|
||
For tickets with non-trivial design questions still open, do another round of targeted
|
||
investigation (fresh `Explore`/`Plan` subagent calls are fine here, run in parallel
|
||
where the questions are independent of each other) until you have concrete, opinionated
|
||
answers — not a menu of options — for every design decision the ticket will hand to an
|
||
implementing subagent.
|
||
|
||
Write one ticket file per ticket into `docs/<slug>/` (pick a short slug from the
|
||
requirements doc's topic — `docs/ui-redesign/`, `docs/feedback/` are the two existing
|
||
precedents in this repo). Write every section in ASD-STE100 (see above). Each ticket
|
||
file uses the same shape every ticket in this repo already uses:
|
||
|
||
```
|
||
# <ID> — <title>
|
||
|
||
**Depends on** <other ticket ids, or —> · **Size** S/M/L · **Status** Not started
|
||
|
||
## Goal
|
||
## Context (exact current code excerpts, file paths, line numbers — quote, don't summarize)
|
||
## Design (concrete, opinionated decisions — not options)
|
||
## Implementation (numbered steps)
|
||
## Acceptance criteria (checkboxes)
|
||
## Tests (specific test names/cases to add, referencing existing test patterns in this repo)
|
||
## Risks
|
||
## Out of scope
|
||
```
|
||
|
||
Also write `docs/<slug>/README.md`: a table of ticket → size → depends-on → status, plus
|
||
a **dependency/dispatch-wave section** (see Phase 4) and a one-paragraph working-method
|
||
summary (implement → analyze/test clean → Outcome section → commit → status flip).
|
||
|
||
Commit the ticket docs as their own commit before dispatching anything.
|
||
|
||
## Phase 4 — Sequence into dispatch waves
|
||
|
||
Two tickets that both edit the same file are a real merge-conflict risk when run in
|
||
parallel isolated worktrees. Group tickets into waves:
|
||
- Tickets with no file overlap with each other → same wave, dispatched in parallel.
|
||
- A ticket that heavily edits a file another ticket already changed → its own later
|
||
wave, dispatched only after the earlier one has merged (so its worktree branches from
|
||
a base that already has the earlier change).
|
||
- Light, disjoint-region overlap in the same file (e.g. two tickets touching different
|
||
methods of the same screen) is an acceptable same-wave risk — it produced a small,
|
||
manually-resolvable conflict in practice, not a blocker.
|
||
|
||
Write this wave plan into the ticket README before dispatching.
|
||
|
||
## Phase 5 — Dispatch, implement, merge (per wave)
|
||
|
||
For each ticket in the current wave, launch an `Agent` call with `isolation: "worktree"`
|
||
and the **full ticket file contents inlined in the prompt** (not just a path — a fresh
|
||
agent should not have to go re-discover context this phase already gathered). Write the
|
||
prompt itself in ASD-STE100 too — the dispatch prompt is the second place, after the
|
||
ticket file, where ambiguity costs the most. The prompt must tell the subagent to:
|
||
|
||
1. Read the inlined ticket, implement exactly what its Design/Implementation sections
|
||
specify — it's already been reviewed, not something to redesign from scratch.
|
||
2. Run `flutter analyze` (must stay clean) and `flutter test` (must stay green, count
|
||
must not regress below the baseline from Phase 0/the previous wave).
|
||
3. Append an `## Outcome` section to its own ticket file: what shipped, any deviation
|
||
from the design and why, the final test count. If an assumption in the ticket turns
|
||
out wrong once real code is touched, say so honestly here rather than silently
|
||
patching over it.
|
||
4. Flip its own ticket's status line to Done — but **not** the shared README's table;
|
||
that gets updated once after merging, by you, not by parallel subagents racing each
|
||
other on the same file.
|
||
5. Commit locally. Follow this repo's normal commit convention unless the user asked
|
||
for something different for this run (e.g. no co-author trailer) — confirm which
|
||
applies during Phase 2 if it's ambiguous, don't assume silently either way.
|
||
6. State explicitly that Android-emulator verification is **best-effort only** at this
|
||
stage (Phase 6 does the real device pass) — a subagent that gets stuck fighting
|
||
emulator flakiness instead of finishing its actual ticket has failed the ticket.
|
||
|
||
Wait for each subagent in the wave to report back, then merge its branch:
|
||
|
||
```bash
|
||
git merge --no-ff worktree-agent-<id> -m "Merge <ticket-id>: <title>"
|
||
```
|
||
|
||
A conflict in the ticket `.md` file itself (both sides changed the status/Outcome
|
||
header) is expected and mechanical — take the incoming side
|
||
(`git checkout --theirs docs/<slug>/<ticket>.md && git add ...`), since it carries the
|
||
Outcome section. A conflict in real code needs an actual read of both diffs — do not
|
||
blindly pick one side; understand what each branch was trying to do and merge the
|
||
*intent*, not just the text (see the worked example: two tickets each added a different,
|
||
correct enhancement to the same function, and the right merge combined both rather than
|
||
picking one).
|
||
|
||
After every merge: `flutter analyze` clean, `flutter test` green, count check. Only then
|
||
move to the next ticket in the wave or the next wave. Update the ticket README's status
|
||
table to Done for everything just merged, as its own small commit.
|
||
|
||
## Phase 6 — Emulator verification pass (all waves merged)
|
||
|
||
This is not optional and not the same thing as the best-effort check individual
|
||
subagents did — this is the real bar. Boot (or reuse, if already running and
|
||
responsive) the project's emulator, launch the freshly-merged app, and walk every
|
||
ticket's golden path for real, taking screenshots as evidence.
|
||
|
||
**Known failure modes from prior runs, and how to handle them without losing an hour:**
|
||
- **Coordinate scaling.** Screenshots are shown to you scaled down (e.g. displayed at
|
||
900×2000 for a native 1080×2400 device) — the image caption states the exact ratio.
|
||
Multiply *both* x and y by that ratio before issuing `adb shell input tap`. Getting
|
||
this wrong looks like "the app isn't responding to taps" and wastes real time — it
|
||
isn't a bug, it's a unit error. When in doubt, pixel-sample the raw screenshot file
|
||
directly (PIL/`Pillow` in Python) to find a real on-screen coordinate rather than
|
||
eyeballing the scaled-down preview.
|
||
- **ANR / "isn't responding" loops.** Check host memory pressure first
|
||
(`vm_stat`, `sysctl vm.swapusage`) before assuming an app bug — a memory-starved host
|
||
produces exactly this symptom. Escalate in order: dismiss-and-wait once → force-stop
|
||
and relaunch the app (`adb shell am force-stop <pkg>` then `am start`) → full emulator
|
||
restart (`adb emu kill` + relaunch) → if the host itself is out of memory, say so and
|
||
ask the user to free some rather than fighting it silently for another hour.
|
||
`com.rippr.port` / `com.rippr.MainActivity` are this app's package/activity for direct
|
||
`am start` (faster than a full `flutter run` rebuild when the APK is already
|
||
installed and you just need a clean process).
|
||
- **Airplane mode / no network** on a fresh emulator boot makes map tiles fail to load
|
||
(renders as a checkerboard, not the app's real tiles) — check
|
||
`adb shell settings get global airplane_mode_on` and
|
||
`adb shell cmd connectivity airplane-mode disable` before concluding tiles are broken.
|
||
- **Stale app state** (a leftover active recording, an old database) from a previous
|
||
manual test session can make a screen look wrong for reasons unrelated to the current
|
||
tickets — `adb shell pm clear <pkg>` before a verification pass if there's any doubt
|
||
the state is clean.
|
||
- **Location-dependent features** (anything reading GPS) need a granted permission and
|
||
a mock fix on the emulator: `adb shell pm grant <pkg> android.permission.ACCESS_FINE_LOCATION`
|
||
(and `ACCESS_COARSE_LOCATION`), then `adb emu geo fix <lon> <lat>` — note longitude
|
||
first. A feature that reads location may take several seconds after granting/fixing
|
||
before the first real fix arrives; don't conclude it's broken from a screenshot taken
|
||
immediately after.
|
||
|
||
**When the emulator pass finds a real bug** (as opposed to an environment artifact
|
||
above): fix it directly, add a regression test that fails without the fix and passes
|
||
with it, re-run `flutter analyze`/`flutter test`, commit with a message that says what
|
||
the emulator pass found and why the widget tests missed it, then re-verify on-device
|
||
that the fix actually shows correctly. Do not mark a ticket's on-device verification
|
||
done until you've seen the fixed behavior with your own eyes (a screenshot), not just
|
||
inferred it from the test passing.
|
||
|
||
If the emulator is genuinely too unstable to get a real signal (host resource
|
||
exhaustion that persists after the escalation steps above), say so plainly rather than
|
||
claiming verification that didn't happen — this exact honesty call was made once
|
||
already in this repo's history and is the correct call to make again if needed.
|
||
|
||
## Phase 7 — Final commit: fresh APK
|
||
|
||
Once every ticket is merged, green, and emulator-verified:
|
||
|
||
```bash
|
||
flutter build apk --release
|
||
git add -f build/app/outputs/flutter-apk/app-release.apk
|
||
git commit -m "<summary of what this batch shipped>"
|
||
```
|
||
|
||
Force-add is required — `build/` is gitignored, and this repo's own established
|
||
precedent (see `git log --oneline -- '*.apk'`) is to force-add the release APK directly
|
||
into the repo rather than relying on CI artifacts or a release page, since there is no
|
||
CI/release pipeline set up. Match that precedent unless the user has since set one up.
|
||
|
||
## Phase 8 — Report back
|
||
|
||
Summarize: what shipped (ticket list with one line each), test count before → after,
|
||
what the emulator pass confirmed (and any bugs it caught that the unit-test bar alone
|
||
would have missed), and the final commit. State plainly whether anything was pushed
|
||
(this skill never pushes on its own — say so, and say what's ready for the user to push
|
||
if they want to).
|
||
|
||
## What this skill is not for
|
||
|
||
A single small, obviously-scoped change (one file, one clear fix) doesn't need five
|
||
phases of ceremony — just make the change, test it, done. This skill is for a
|
||
requirements document with enough surface area that breaking it into independently
|
||
verifiable pieces is itself valuable — the kind of input that would otherwise become one
|
||
enormous, hard-to-review commit.
|