Add autopilot-build skill: requirements-to-shipped-APK workflow

Codifies the ticket-writing, human-validation, wave-dispatch,
worktree-merge, emulator-verification, and APK-commit loop this repo's
UI-redesign and feedback batches were built with, written in ASD-STE100
so ticket files and subagent prompts stay unambiguous for a fresh
subagent with no conversation history.
This commit is contained in:
2026-08-24 16:50:58 -05:00
parent f30f0063d0
commit e06b371e2d
2 changed files with 271 additions and 1 deletions

View File

@@ -0,0 +1,270 @@
---
name: autopilot-build
description: "Turn a markdown requirements file into implemented, tested, emulator-verified, committed Flutter changes for this repo (rippr-flutter) — breaks requirements into independently-completable tickets, pauses for human approval, then dispatches implementation to isolated subagents, merges, verifies on the Android emulator, and commits a fresh APK."
---
# Autopilot Build
Runs the same disciplined loop this repo's UI-redesign (`docs/ui-redesign/`) and
post-launch-feedback (`docs/feedback/`) batches were built with, end to end, from a
single requirements document. It is slow and thorough by design — every phase below
exists because skipping it caused a real problem in a prior run of this loop.
## Writing style: ASD-STE100 (Simplified Technical English)
Every ticket file, `Outcome` section, and subagent dispatch prompt this skill produces
is written in **ASD-STE100** — the controlled-language standard aerospace maintenance
manuals use specifically to remove ambiguity for a reader with no shared context. That
is exactly the constraint a fresh subagent with zero conversation history is under, so
the standard's core rules apply directly, not as a style nicety:
- **One instruction, one sentence.** Never join two instructions with "and" or a comma.
"Remove the panel. Disconnect the cable." not "Remove the panel and disconnect the
cable."
- **Active voice, imperative for instructions.** "Add the constant to `ride_map.dart`."
not "The constant should be added to `ride_map.dart`."
- **Keep sentences short** (roughly 20 words or fewer). Split a long sentence at its
natural joints rather than adding a subordinate clause.
- **Use one word for one meaning, every time.** If a ticket calls it "the overflow
menu" in one place, never call the same thing "the options menu" or "the kebab menu"
later in the same ticket. Inconsistent naming is the single most common way a
fresh-context reader misidentifies which code you mean.
- **Avoid long noun strings.** "the HUD widget grid collision check" is four nouns
stacked — prefer "the check that finds a collision in the HUD widget grid."
- **Avoid ambiguous pronouns.** Repeat the noun instead of "it"/"this"/"that" whenever
more than one candidate referent exists in the surrounding sentences — the exact
failure mode a fresh subagent can't recover from is guessing which prior noun "it"
meant.
- **Prefer concrete, specific verbs over vague ones.** "Reject the move" not "handle
the move"; "return the cached value" not "deal with the cached value."
- **State conditions before actions.** "If the route has fewer than two waypoints,
disable the toggle." not "Disable the toggle if the route has fewer than two
waypoints" — STE100 puts the condition first so the reader isn't committed to an
action before learning it might not apply.
This does not apply to direct quotations of existing code or user feedback (quote those
verbatim, exactly as written) — only to the ticket prose you write around them.
## Input
One argument: a path to a markdown requirements file (e.g. `docs/FEEDBACK.md`, or
whatever the user hands you). If no path is given, ask for one — don't guess at
requirements from conversation context alone; the file is the source of truth this
whole run is grounded in.
## Phase 0 — Sanity check before starting
- `flutter analyze` clean and `flutter test` green. Record the passing count — every
later phase's bar is "still clean, count only goes up from here."
- `git status` clean (or confirm with the user what to do with any dirty state — don't
silently work on top of someone else's uncommitted changes).
- Confirm a remote exists if this run is expected to end in something pushable — this
skill never pushes itself (see Phase 8), but say so up front if there's no remote.
## Phase 1 — Break requirements into tickets
Read the requirements file in full. Read enough of the current codebase (grep, targeted
file reads, and/or `Explore`-type subagents if the surface area is large or unfamiliar)
to ground each requirement in real files, not guesses. For anything with real design
ambiguity (a new data model, a rendering bug with no obvious cause, an architecture
change), launch a `Plan`-type subagent to work through the approach *before* you commit
to a ticket list — cheaper to redesign a paragraph now than a shipped ticket later.
Split the requirements into a small number of tickets, each one:
- **Independently completable** — a fresh subagent with zero conversation history could
read the ticket file alone and implement it correctly. This is the single most
important property. If a ticket needs context only available in this conversation,
the ticket is incomplete, not the subagent.
- **Scoped to avoid unnecessary file overlap** with its siblings where possible. Where
overlap is unavoidable (two tickets both need to touch the same screen), note the
dependency explicitly — see Phase 2.
**Do not write full ticket files yet.** Recite the list to the user first: one line
per ticket (title, one-sentence scope, rough size S/M/L). This mirrors the real
worked example: five tickets were named and sized in one message, and only written out
in full after the user confirmed the list. Skipping this step means discovering scope
disagreements after hours of writing detail nobody asked for.
## Phase 2 — Human validation (hard stop)
Do not proceed to Phase 3 without an explicit go-ahead on the ticket list. Use
`AskUserQuestion` if there's a real fork to resolve (e.g. "should FB-04 and FB-05 be one
ticket or two?"), otherwise just present the recited list and wait for the user's next
message. Adjust scope, splitting, or sizing based on their response before moving on —
this is the point where it's cheap to change direction.
## Phase 3 — Deep-dive each ticket and write full ticket files
For tickets with non-trivial design questions still open, do another round of targeted
investigation (fresh `Explore`/`Plan` subagent calls are fine here, run in parallel
where the questions are independent of each other) until you have concrete, opinionated
answers — not a menu of options — for every design decision the ticket will hand to an
implementing subagent.
Write one ticket file per ticket into `docs/<slug>/` (pick a short slug from the
requirements doc's topic — `docs/ui-redesign/`, `docs/feedback/` are the two existing
precedents in this repo). Write every section in ASD-STE100 (see above). Each ticket
file uses the same shape every ticket in this repo already uses:
```
# <ID> — <title>
**Depends on** <other ticket ids, or —> · **Size** S/M/L · **Status** Not started
## Goal
## Context (exact current code excerpts, file paths, line numbers — quote, don't summarize)
## Design (concrete, opinionated decisions — not options)
## Implementation (numbered steps)
## Acceptance criteria (checkboxes)
## Tests (specific test names/cases to add, referencing existing test patterns in this repo)
## Risks
## Out of scope
```
Also write `docs/<slug>/README.md`: a table of ticket → size → depends-on → status, plus
a **dependency/dispatch-wave section** (see Phase 4) and a one-paragraph working-method
summary (implement → analyze/test clean → Outcome section → commit → status flip).
Commit the ticket docs as their own commit before dispatching anything.
## Phase 4 — Sequence into dispatch waves
Two tickets that both edit the same file are a real merge-conflict risk when run in
parallel isolated worktrees. Group tickets into waves:
- Tickets with no file overlap with each other → same wave, dispatched in parallel.
- A ticket that heavily edits a file another ticket already changed → its own later
wave, dispatched only after the earlier one has merged (so its worktree branches from
a base that already has the earlier change).
- Light, disjoint-region overlap in the same file (e.g. two tickets touching different
methods of the same screen) is an acceptable same-wave risk — it produced a small,
manually-resolvable conflict in practice, not a blocker.
Write this wave plan into the ticket README before dispatching.
## Phase 5 — Dispatch, implement, merge (per wave)
For each ticket in the current wave, launch an `Agent` call with `isolation: "worktree"`
and the **full ticket file contents inlined in the prompt** (not just a path — a fresh
agent should not have to go re-discover context this phase already gathered). Write the
prompt itself in ASD-STE100 too — the dispatch prompt is the second place, after the
ticket file, where ambiguity costs the most. The prompt must tell the subagent to:
1. Read the inlined ticket, implement exactly what its Design/Implementation sections
specify — it's already been reviewed, not something to redesign from scratch.
2. Run `flutter analyze` (must stay clean) and `flutter test` (must stay green, count
must not regress below the baseline from Phase 0/the previous wave).
3. Append an `## Outcome` section to its own ticket file: what shipped, any deviation
from the design and why, the final test count. If an assumption in the ticket turns
out wrong once real code is touched, say so honestly here rather than silently
patching over it.
4. Flip its own ticket's status line to Done — but **not** the shared README's table;
that gets updated once after merging, by you, not by parallel subagents racing each
other on the same file.
5. Commit locally. Follow this repo's normal commit convention unless the user asked
for something different for this run (e.g. no co-author trailer) — confirm which
applies during Phase 2 if it's ambiguous, don't assume silently either way.
6. State explicitly that Android-emulator verification is **best-effort only** at this
stage (Phase 6 does the real device pass) — a subagent that gets stuck fighting
emulator flakiness instead of finishing its actual ticket has failed the ticket.
Wait for each subagent in the wave to report back, then merge its branch:
```bash
git merge --no-ff worktree-agent-<id> -m "Merge <ticket-id>: <title>"
```
A conflict in the ticket `.md` file itself (both sides changed the status/Outcome
header) is expected and mechanical — take the incoming side
(`git checkout --theirs docs/<slug>/<ticket>.md && git add ...`), since it carries the
Outcome section. A conflict in real code needs an actual read of both diffs — do not
blindly pick one side; understand what each branch was trying to do and merge the
*intent*, not just the text (see the worked example: two tickets each added a different,
correct enhancement to the same function, and the right merge combined both rather than
picking one).
After every merge: `flutter analyze` clean, `flutter test` green, count check. Only then
move to the next ticket in the wave or the next wave. Update the ticket README's status
table to Done for everything just merged, as its own small commit.
## Phase 6 — Emulator verification pass (all waves merged)
This is not optional and not the same thing as the best-effort check individual
subagents did — this is the real bar. Boot (or reuse, if already running and
responsive) the project's emulator, launch the freshly-merged app, and walk every
ticket's golden path for real, taking screenshots as evidence.
**Known failure modes from prior runs, and how to handle them without losing an hour:**
- **Coordinate scaling.** Screenshots are shown to you scaled down (e.g. displayed at
900×2000 for a native 1080×2400 device) — the image caption states the exact ratio.
Multiply *both* x and y by that ratio before issuing `adb shell input tap`. Getting
this wrong looks like "the app isn't responding to taps" and wastes real time — it
isn't a bug, it's a unit error. When in doubt, pixel-sample the raw screenshot file
directly (PIL/`Pillow` in Python) to find a real on-screen coordinate rather than
eyeballing the scaled-down preview.
- **ANR / "isn't responding" loops.** Check host memory pressure first
(`vm_stat`, `sysctl vm.swapusage`) before assuming an app bug — a memory-starved host
produces exactly this symptom. Escalate in order: dismiss-and-wait once → force-stop
and relaunch the app (`adb shell am force-stop <pkg>` then `am start`) → full emulator
restart (`adb emu kill` + relaunch) → if the host itself is out of memory, say so and
ask the user to free some rather than fighting it silently for another hour.
`com.rippr.port` / `com.rippr.MainActivity` are this app's package/activity for direct
`am start` (faster than a full `flutter run` rebuild when the APK is already
installed and you just need a clean process).
- **Airplane mode / no network** on a fresh emulator boot makes map tiles fail to load
(renders as a checkerboard, not the app's real tiles) — check
`adb shell settings get global airplane_mode_on` and
`adb shell cmd connectivity airplane-mode disable` before concluding tiles are broken.
- **Stale app state** (a leftover active recording, an old database) from a previous
manual test session can make a screen look wrong for reasons unrelated to the current
tickets — `adb shell pm clear <pkg>` before a verification pass if there's any doubt
the state is clean.
- **Location-dependent features** (anything reading GPS) need a granted permission and
a mock fix on the emulator: `adb shell pm grant <pkg> android.permission.ACCESS_FINE_LOCATION`
(and `ACCESS_COARSE_LOCATION`), then `adb emu geo fix <lon> <lat>` — note longitude
first. A feature that reads location may take several seconds after granting/fixing
before the first real fix arrives; don't conclude it's broken from a screenshot taken
immediately after.
**When the emulator pass finds a real bug** (as opposed to an environment artifact
above): fix it directly, add a regression test that fails without the fix and passes
with it, re-run `flutter analyze`/`flutter test`, commit with a message that says what
the emulator pass found and why the widget tests missed it, then re-verify on-device
that the fix actually shows correctly. Do not mark a ticket's on-device verification
done until you've seen the fixed behavior with your own eyes (a screenshot), not just
inferred it from the test passing.
If the emulator is genuinely too unstable to get a real signal (host resource
exhaustion that persists after the escalation steps above), say so plainly rather than
claiming verification that didn't happen — this exact honesty call was made once
already in this repo's history and is the correct call to make again if needed.
## Phase 7 — Final commit: fresh APK
Once every ticket is merged, green, and emulator-verified:
```bash
flutter build apk --release
git add -f build/app/outputs/flutter-apk/app-release.apk
git commit -m "<summary of what this batch shipped>"
```
Force-add is required — `build/` is gitignored, and this repo's own established
precedent (see `git log --oneline -- '*.apk'`) is to force-add the release APK directly
into the repo rather than relying on CI artifacts or a release page, since there is no
CI/release pipeline set up. Match that precedent unless the user has since set one up.
## Phase 8 — Report back
Summarize: what shipped (ticket list with one line each), test count before → after,
what the emulator pass confirmed (and any bugs it caught that the unit-test bar alone
would have missed), and the final commit. State plainly whether anything was pushed
(this skill never pushes on its own — say so, and say what's ready for the user to push
if they want to).
## What this skill is not for
A single small, obviously-scoped change (one file, one clear fix) doesn't need five
phases of ceremony — just make the change, test it, done. This skill is for a
requirements document with enough surface area that breaking it into independently
verifiable pieces is itself valuable — the kind of input that would otherwise become one
enormous, hard-to-review commit.

2
.gitignore vendored
View File

@@ -50,4 +50,4 @@ app.*.map.json
# Build artifacts (large; regenerable)
build/
.dart_tool/
.claude/
.claude/worktrees/