Add autopilot-build skill: requirements-to-shipped-APK workflow
Codifies the ticket-writing, human-validation, wave-dispatch, worktree-merge, emulator-verification, and APK-commit loop this repo's UI-redesign and feedback batches were built with, written in ASD-STE100 so ticket files and subagent prompts stay unambiguous for a fresh subagent with no conversation history.
This commit is contained in:
270
.claude/skills/autopilot-build/SKILL.md
Normal file
270
.claude/skills/autopilot-build/SKILL.md
Normal file
@@ -0,0 +1,270 @@
|
||||
---
|
||||
name: autopilot-build
|
||||
description: "Turn a markdown requirements file into implemented, tested, emulator-verified, committed Flutter changes for this repo (rippr-flutter) — breaks requirements into independently-completable tickets, pauses for human approval, then dispatches implementation to isolated subagents, merges, verifies on the Android emulator, and commits a fresh APK."
|
||||
---
|
||||
|
||||
# Autopilot Build
|
||||
|
||||
Runs the same disciplined loop this repo's UI-redesign (`docs/ui-redesign/`) and
|
||||
post-launch-feedback (`docs/feedback/`) batches were built with, end to end, from a
|
||||
single requirements document. It is slow and thorough by design — every phase below
|
||||
exists because skipping it caused a real problem in a prior run of this loop.
|
||||
|
||||
## Writing style: ASD-STE100 (Simplified Technical English)
|
||||
|
||||
Every ticket file, `Outcome` section, and subagent dispatch prompt this skill produces
|
||||
is written in **ASD-STE100** — the controlled-language standard aerospace maintenance
|
||||
manuals use specifically to remove ambiguity for a reader with no shared context. That
|
||||
is exactly the constraint a fresh subagent with zero conversation history is under, so
|
||||
the standard's core rules apply directly, not as a style nicety:
|
||||
|
||||
- **One instruction, one sentence.** Never join two instructions with "and" or a comma.
|
||||
"Remove the panel. Disconnect the cable." not "Remove the panel and disconnect the
|
||||
cable."
|
||||
- **Active voice, imperative for instructions.** "Add the constant to `ride_map.dart`."
|
||||
not "The constant should be added to `ride_map.dart`."
|
||||
- **Keep sentences short** (roughly 20 words or fewer). Split a long sentence at its
|
||||
natural joints rather than adding a subordinate clause.
|
||||
- **Use one word for one meaning, every time.** If a ticket calls it "the overflow
|
||||
menu" in one place, never call the same thing "the options menu" or "the kebab menu"
|
||||
later in the same ticket. Inconsistent naming is the single most common way a
|
||||
fresh-context reader misidentifies which code you mean.
|
||||
- **Avoid long noun strings.** "the HUD widget grid collision check" is four nouns
|
||||
stacked — prefer "the check that finds a collision in the HUD widget grid."
|
||||
- **Avoid ambiguous pronouns.** Repeat the noun instead of "it"/"this"/"that" whenever
|
||||
more than one candidate referent exists in the surrounding sentences — the exact
|
||||
failure mode a fresh subagent can't recover from is guessing which prior noun "it"
|
||||
meant.
|
||||
- **Prefer concrete, specific verbs over vague ones.** "Reject the move" not "handle
|
||||
the move"; "return the cached value" not "deal with the cached value."
|
||||
- **State conditions before actions.** "If the route has fewer than two waypoints,
|
||||
disable the toggle." not "Disable the toggle if the route has fewer than two
|
||||
waypoints" — STE100 puts the condition first so the reader isn't committed to an
|
||||
action before learning it might not apply.
|
||||
|
||||
This does not apply to direct quotations of existing code or user feedback (quote those
|
||||
verbatim, exactly as written) — only to the ticket prose you write around them.
|
||||
|
||||
## Input
|
||||
|
||||
One argument: a path to a markdown requirements file (e.g. `docs/FEEDBACK.md`, or
|
||||
whatever the user hands you). If no path is given, ask for one — don't guess at
|
||||
requirements from conversation context alone; the file is the source of truth this
|
||||
whole run is grounded in.
|
||||
|
||||
## Phase 0 — Sanity check before starting
|
||||
|
||||
- `flutter analyze` clean and `flutter test` green. Record the passing count — every
|
||||
later phase's bar is "still clean, count only goes up from here."
|
||||
- `git status` clean (or confirm with the user what to do with any dirty state — don't
|
||||
silently work on top of someone else's uncommitted changes).
|
||||
- Confirm a remote exists if this run is expected to end in something pushable — this
|
||||
skill never pushes itself (see Phase 8), but say so up front if there's no remote.
|
||||
|
||||
## Phase 1 — Break requirements into tickets
|
||||
|
||||
Read the requirements file in full. Read enough of the current codebase (grep, targeted
|
||||
file reads, and/or `Explore`-type subagents if the surface area is large or unfamiliar)
|
||||
to ground each requirement in real files, not guesses. For anything with real design
|
||||
ambiguity (a new data model, a rendering bug with no obvious cause, an architecture
|
||||
change), launch a `Plan`-type subagent to work through the approach *before* you commit
|
||||
to a ticket list — cheaper to redesign a paragraph now than a shipped ticket later.
|
||||
|
||||
Split the requirements into a small number of tickets, each one:
|
||||
- **Independently completable** — a fresh subagent with zero conversation history could
|
||||
read the ticket file alone and implement it correctly. This is the single most
|
||||
important property. If a ticket needs context only available in this conversation,
|
||||
the ticket is incomplete, not the subagent.
|
||||
- **Scoped to avoid unnecessary file overlap** with its siblings where possible. Where
|
||||
overlap is unavoidable (two tickets both need to touch the same screen), note the
|
||||
dependency explicitly — see Phase 2.
|
||||
|
||||
**Do not write full ticket files yet.** Recite the list to the user first: one line
|
||||
per ticket (title, one-sentence scope, rough size S/M/L). This mirrors the real
|
||||
worked example: five tickets were named and sized in one message, and only written out
|
||||
in full after the user confirmed the list. Skipping this step means discovering scope
|
||||
disagreements after hours of writing detail nobody asked for.
|
||||
|
||||
## Phase 2 — Human validation (hard stop)
|
||||
|
||||
Do not proceed to Phase 3 without an explicit go-ahead on the ticket list. Use
|
||||
`AskUserQuestion` if there's a real fork to resolve (e.g. "should FB-04 and FB-05 be one
|
||||
ticket or two?"), otherwise just present the recited list and wait for the user's next
|
||||
message. Adjust scope, splitting, or sizing based on their response before moving on —
|
||||
this is the point where it's cheap to change direction.
|
||||
|
||||
## Phase 3 — Deep-dive each ticket and write full ticket files
|
||||
|
||||
For tickets with non-trivial design questions still open, do another round of targeted
|
||||
investigation (fresh `Explore`/`Plan` subagent calls are fine here, run in parallel
|
||||
where the questions are independent of each other) until you have concrete, opinionated
|
||||
answers — not a menu of options — for every design decision the ticket will hand to an
|
||||
implementing subagent.
|
||||
|
||||
Write one ticket file per ticket into `docs/<slug>/` (pick a short slug from the
|
||||
requirements doc's topic — `docs/ui-redesign/`, `docs/feedback/` are the two existing
|
||||
precedents in this repo). Write every section in ASD-STE100 (see above). Each ticket
|
||||
file uses the same shape every ticket in this repo already uses:
|
||||
|
||||
```
|
||||
# <ID> — <title>
|
||||
|
||||
**Depends on** <other ticket ids, or —> · **Size** S/M/L · **Status** Not started
|
||||
|
||||
## Goal
|
||||
## Context (exact current code excerpts, file paths, line numbers — quote, don't summarize)
|
||||
## Design (concrete, opinionated decisions — not options)
|
||||
## Implementation (numbered steps)
|
||||
## Acceptance criteria (checkboxes)
|
||||
## Tests (specific test names/cases to add, referencing existing test patterns in this repo)
|
||||
## Risks
|
||||
## Out of scope
|
||||
```
|
||||
|
||||
Also write `docs/<slug>/README.md`: a table of ticket → size → depends-on → status, plus
|
||||
a **dependency/dispatch-wave section** (see Phase 4) and a one-paragraph working-method
|
||||
summary (implement → analyze/test clean → Outcome section → commit → status flip).
|
||||
|
||||
Commit the ticket docs as their own commit before dispatching anything.
|
||||
|
||||
## Phase 4 — Sequence into dispatch waves
|
||||
|
||||
Two tickets that both edit the same file are a real merge-conflict risk when run in
|
||||
parallel isolated worktrees. Group tickets into waves:
|
||||
- Tickets with no file overlap with each other → same wave, dispatched in parallel.
|
||||
- A ticket that heavily edits a file another ticket already changed → its own later
|
||||
wave, dispatched only after the earlier one has merged (so its worktree branches from
|
||||
a base that already has the earlier change).
|
||||
- Light, disjoint-region overlap in the same file (e.g. two tickets touching different
|
||||
methods of the same screen) is an acceptable same-wave risk — it produced a small,
|
||||
manually-resolvable conflict in practice, not a blocker.
|
||||
|
||||
Write this wave plan into the ticket README before dispatching.
|
||||
|
||||
## Phase 5 — Dispatch, implement, merge (per wave)
|
||||
|
||||
For each ticket in the current wave, launch an `Agent` call with `isolation: "worktree"`
|
||||
and the **full ticket file contents inlined in the prompt** (not just a path — a fresh
|
||||
agent should not have to go re-discover context this phase already gathered). Write the
|
||||
prompt itself in ASD-STE100 too — the dispatch prompt is the second place, after the
|
||||
ticket file, where ambiguity costs the most. The prompt must tell the subagent to:
|
||||
|
||||
1. Read the inlined ticket, implement exactly what its Design/Implementation sections
|
||||
specify — it's already been reviewed, not something to redesign from scratch.
|
||||
2. Run `flutter analyze` (must stay clean) and `flutter test` (must stay green, count
|
||||
must not regress below the baseline from Phase 0/the previous wave).
|
||||
3. Append an `## Outcome` section to its own ticket file: what shipped, any deviation
|
||||
from the design and why, the final test count. If an assumption in the ticket turns
|
||||
out wrong once real code is touched, say so honestly here rather than silently
|
||||
patching over it.
|
||||
4. Flip its own ticket's status line to Done — but **not** the shared README's table;
|
||||
that gets updated once after merging, by you, not by parallel subagents racing each
|
||||
other on the same file.
|
||||
5. Commit locally. Follow this repo's normal commit convention unless the user asked
|
||||
for something different for this run (e.g. no co-author trailer) — confirm which
|
||||
applies during Phase 2 if it's ambiguous, don't assume silently either way.
|
||||
6. State explicitly that Android-emulator verification is **best-effort only** at this
|
||||
stage (Phase 6 does the real device pass) — a subagent that gets stuck fighting
|
||||
emulator flakiness instead of finishing its actual ticket has failed the ticket.
|
||||
|
||||
Wait for each subagent in the wave to report back, then merge its branch:
|
||||
|
||||
```bash
|
||||
git merge --no-ff worktree-agent-<id> -m "Merge <ticket-id>: <title>"
|
||||
```
|
||||
|
||||
A conflict in the ticket `.md` file itself (both sides changed the status/Outcome
|
||||
header) is expected and mechanical — take the incoming side
|
||||
(`git checkout --theirs docs/<slug>/<ticket>.md && git add ...`), since it carries the
|
||||
Outcome section. A conflict in real code needs an actual read of both diffs — do not
|
||||
blindly pick one side; understand what each branch was trying to do and merge the
|
||||
*intent*, not just the text (see the worked example: two tickets each added a different,
|
||||
correct enhancement to the same function, and the right merge combined both rather than
|
||||
picking one).
|
||||
|
||||
After every merge: `flutter analyze` clean, `flutter test` green, count check. Only then
|
||||
move to the next ticket in the wave or the next wave. Update the ticket README's status
|
||||
table to Done for everything just merged, as its own small commit.
|
||||
|
||||
## Phase 6 — Emulator verification pass (all waves merged)
|
||||
|
||||
This is not optional and not the same thing as the best-effort check individual
|
||||
subagents did — this is the real bar. Boot (or reuse, if already running and
|
||||
responsive) the project's emulator, launch the freshly-merged app, and walk every
|
||||
ticket's golden path for real, taking screenshots as evidence.
|
||||
|
||||
**Known failure modes from prior runs, and how to handle them without losing an hour:**
|
||||
- **Coordinate scaling.** Screenshots are shown to you scaled down (e.g. displayed at
|
||||
900×2000 for a native 1080×2400 device) — the image caption states the exact ratio.
|
||||
Multiply *both* x and y by that ratio before issuing `adb shell input tap`. Getting
|
||||
this wrong looks like "the app isn't responding to taps" and wastes real time — it
|
||||
isn't a bug, it's a unit error. When in doubt, pixel-sample the raw screenshot file
|
||||
directly (PIL/`Pillow` in Python) to find a real on-screen coordinate rather than
|
||||
eyeballing the scaled-down preview.
|
||||
- **ANR / "isn't responding" loops.** Check host memory pressure first
|
||||
(`vm_stat`, `sysctl vm.swapusage`) before assuming an app bug — a memory-starved host
|
||||
produces exactly this symptom. Escalate in order: dismiss-and-wait once → force-stop
|
||||
and relaunch the app (`adb shell am force-stop <pkg>` then `am start`) → full emulator
|
||||
restart (`adb emu kill` + relaunch) → if the host itself is out of memory, say so and
|
||||
ask the user to free some rather than fighting it silently for another hour.
|
||||
`com.rippr.port` / `com.rippr.MainActivity` are this app's package/activity for direct
|
||||
`am start` (faster than a full `flutter run` rebuild when the APK is already
|
||||
installed and you just need a clean process).
|
||||
- **Airplane mode / no network** on a fresh emulator boot makes map tiles fail to load
|
||||
(renders as a checkerboard, not the app's real tiles) — check
|
||||
`adb shell settings get global airplane_mode_on` and
|
||||
`adb shell cmd connectivity airplane-mode disable` before concluding tiles are broken.
|
||||
- **Stale app state** (a leftover active recording, an old database) from a previous
|
||||
manual test session can make a screen look wrong for reasons unrelated to the current
|
||||
tickets — `adb shell pm clear <pkg>` before a verification pass if there's any doubt
|
||||
the state is clean.
|
||||
- **Location-dependent features** (anything reading GPS) need a granted permission and
|
||||
a mock fix on the emulator: `adb shell pm grant <pkg> android.permission.ACCESS_FINE_LOCATION`
|
||||
(and `ACCESS_COARSE_LOCATION`), then `adb emu geo fix <lon> <lat>` — note longitude
|
||||
first. A feature that reads location may take several seconds after granting/fixing
|
||||
before the first real fix arrives; don't conclude it's broken from a screenshot taken
|
||||
immediately after.
|
||||
|
||||
**When the emulator pass finds a real bug** (as opposed to an environment artifact
|
||||
above): fix it directly, add a regression test that fails without the fix and passes
|
||||
with it, re-run `flutter analyze`/`flutter test`, commit with a message that says what
|
||||
the emulator pass found and why the widget tests missed it, then re-verify on-device
|
||||
that the fix actually shows correctly. Do not mark a ticket's on-device verification
|
||||
done until you've seen the fixed behavior with your own eyes (a screenshot), not just
|
||||
inferred it from the test passing.
|
||||
|
||||
If the emulator is genuinely too unstable to get a real signal (host resource
|
||||
exhaustion that persists after the escalation steps above), say so plainly rather than
|
||||
claiming verification that didn't happen — this exact honesty call was made once
|
||||
already in this repo's history and is the correct call to make again if needed.
|
||||
|
||||
## Phase 7 — Final commit: fresh APK
|
||||
|
||||
Once every ticket is merged, green, and emulator-verified:
|
||||
|
||||
```bash
|
||||
flutter build apk --release
|
||||
git add -f build/app/outputs/flutter-apk/app-release.apk
|
||||
git commit -m "<summary of what this batch shipped>"
|
||||
```
|
||||
|
||||
Force-add is required — `build/` is gitignored, and this repo's own established
|
||||
precedent (see `git log --oneline -- '*.apk'`) is to force-add the release APK directly
|
||||
into the repo rather than relying on CI artifacts or a release page, since there is no
|
||||
CI/release pipeline set up. Match that precedent unless the user has since set one up.
|
||||
|
||||
## Phase 8 — Report back
|
||||
|
||||
Summarize: what shipped (ticket list with one line each), test count before → after,
|
||||
what the emulator pass confirmed (and any bugs it caught that the unit-test bar alone
|
||||
would have missed), and the final commit. State plainly whether anything was pushed
|
||||
(this skill never pushes on its own — say so, and say what's ready for the user to push
|
||||
if they want to).
|
||||
|
||||
## What this skill is not for
|
||||
|
||||
A single small, obviously-scoped change (one file, one clear fix) doesn't need five
|
||||
phases of ceremony — just make the change, test it, done. This skill is for a
|
||||
requirements document with enough surface area that breaking it into independently
|
||||
verifiable pieces is itself valuable — the kind of input that would otherwise become one
|
||||
enormous, hard-to-review commit.
|
||||
2
.gitignore
vendored
2
.gitignore
vendored
@@ -50,4 +50,4 @@ app.*.map.json
|
||||
# Build artifacts (large; regenerable)
|
||||
build/
|
||||
.dart_tool/
|
||||
.claude/
|
||||
.claude/worktrees/
|
||||
|
||||
Reference in New Issue
Block a user