diff --git a/.claude/skills/autopilot-build/SKILL.md b/.claude/skills/autopilot-build/SKILL.md new file mode 100644 index 0000000..a946104 --- /dev/null +++ b/.claude/skills/autopilot-build/SKILL.md @@ -0,0 +1,270 @@ +--- +name: autopilot-build +description: "Turn a markdown requirements file into implemented, tested, emulator-verified, committed Flutter changes for this repo (rippr-flutter) — breaks requirements into independently-completable tickets, pauses for human approval, then dispatches implementation to isolated subagents, merges, verifies on the Android emulator, and commits a fresh APK." +--- + +# Autopilot Build + +Runs the same disciplined loop this repo's UI-redesign (`docs/ui-redesign/`) and +post-launch-feedback (`docs/feedback/`) batches were built with, end to end, from a +single requirements document. It is slow and thorough by design — every phase below +exists because skipping it caused a real problem in a prior run of this loop. + +## Writing style: ASD-STE100 (Simplified Technical English) + +Every ticket file, `Outcome` section, and subagent dispatch prompt this skill produces +is written in **ASD-STE100** — the controlled-language standard aerospace maintenance +manuals use specifically to remove ambiguity for a reader with no shared context. That +is exactly the constraint a fresh subagent with zero conversation history is under, so +the standard's core rules apply directly, not as a style nicety: + +- **One instruction, one sentence.** Never join two instructions with "and" or a comma. + "Remove the panel. Disconnect the cable." not "Remove the panel and disconnect the + cable." +- **Active voice, imperative for instructions.** "Add the constant to `ride_map.dart`." + not "The constant should be added to `ride_map.dart`." +- **Keep sentences short** (roughly 20 words or fewer). Split a long sentence at its + natural joints rather than adding a subordinate clause. +- **Use one word for one meaning, every time.** If a ticket calls it "the overflow + menu" in one place, never call the same thing "the options menu" or "the kebab menu" + later in the same ticket. Inconsistent naming is the single most common way a + fresh-context reader misidentifies which code you mean. +- **Avoid long noun strings.** "the HUD widget grid collision check" is four nouns + stacked — prefer "the check that finds a collision in the HUD widget grid." +- **Avoid ambiguous pronouns.** Repeat the noun instead of "it"/"this"/"that" whenever + more than one candidate referent exists in the surrounding sentences — the exact + failure mode a fresh subagent can't recover from is guessing which prior noun "it" + meant. +- **Prefer concrete, specific verbs over vague ones.** "Reject the move" not "handle + the move"; "return the cached value" not "deal with the cached value." +- **State conditions before actions.** "If the route has fewer than two waypoints, + disable the toggle." not "Disable the toggle if the route has fewer than two + waypoints" — STE100 puts the condition first so the reader isn't committed to an + action before learning it might not apply. + +This does not apply to direct quotations of existing code or user feedback (quote those +verbatim, exactly as written) — only to the ticket prose you write around them. + +## Input + +One argument: a path to a markdown requirements file (e.g. `docs/FEEDBACK.md`, or +whatever the user hands you). If no path is given, ask for one — don't guess at +requirements from conversation context alone; the file is the source of truth this +whole run is grounded in. + +## Phase 0 — Sanity check before starting + +- `flutter analyze` clean and `flutter test` green. Record the passing count — every + later phase's bar is "still clean, count only goes up from here." +- `git status` clean (or confirm with the user what to do with any dirty state — don't + silently work on top of someone else's uncommitted changes). +- Confirm a remote exists if this run is expected to end in something pushable — this + skill never pushes itself (see Phase 8), but say so up front if there's no remote. + +## Phase 1 — Break requirements into tickets + +Read the requirements file in full. Read enough of the current codebase (grep, targeted +file reads, and/or `Explore`-type subagents if the surface area is large or unfamiliar) +to ground each requirement in real files, not guesses. For anything with real design +ambiguity (a new data model, a rendering bug with no obvious cause, an architecture +change), launch a `Plan`-type subagent to work through the approach *before* you commit +to a ticket list — cheaper to redesign a paragraph now than a shipped ticket later. + +Split the requirements into a small number of tickets, each one: +- **Independently completable** — a fresh subagent with zero conversation history could + read the ticket file alone and implement it correctly. This is the single most + important property. If a ticket needs context only available in this conversation, + the ticket is incomplete, not the subagent. +- **Scoped to avoid unnecessary file overlap** with its siblings where possible. Where + overlap is unavoidable (two tickets both need to touch the same screen), note the + dependency explicitly — see Phase 2. + +**Do not write full ticket files yet.** Recite the list to the user first: one line +per ticket (title, one-sentence scope, rough size S/M/L). This mirrors the real +worked example: five tickets were named and sized in one message, and only written out +in full after the user confirmed the list. Skipping this step means discovering scope +disagreements after hours of writing detail nobody asked for. + +## Phase 2 — Human validation (hard stop) + +Do not proceed to Phase 3 without an explicit go-ahead on the ticket list. Use +`AskUserQuestion` if there's a real fork to resolve (e.g. "should FB-04 and FB-05 be one +ticket or two?"), otherwise just present the recited list and wait for the user's next +message. Adjust scope, splitting, or sizing based on their response before moving on — +this is the point where it's cheap to change direction. + +## Phase 3 — Deep-dive each ticket and write full ticket files + +For tickets with non-trivial design questions still open, do another round of targeted +investigation (fresh `Explore`/`Plan` subagent calls are fine here, run in parallel +where the questions are independent of each other) until you have concrete, opinionated +answers — not a menu of options — for every design decision the ticket will hand to an +implementing subagent. + +Write one ticket file per ticket into `docs//` (pick a short slug from the +requirements doc's topic — `docs/ui-redesign/`, `docs/feedback/` are the two existing +precedents in this repo). Write every section in ASD-STE100 (see above). Each ticket +file uses the same shape every ticket in this repo already uses: + +``` +# — + +**Depends on** <other ticket ids, or —> · **Size** S/M/L · **Status** Not started + +## Goal +## Context (exact current code excerpts, file paths, line numbers — quote, don't summarize) +## Design (concrete, opinionated decisions — not options) +## Implementation (numbered steps) +## Acceptance criteria (checkboxes) +## Tests (specific test names/cases to add, referencing existing test patterns in this repo) +## Risks +## Out of scope +``` + +Also write `docs/<slug>/README.md`: a table of ticket → size → depends-on → status, plus +a **dependency/dispatch-wave section** (see Phase 4) and a one-paragraph working-method +summary (implement → analyze/test clean → Outcome section → commit → status flip). + +Commit the ticket docs as their own commit before dispatching anything. + +## Phase 4 — Sequence into dispatch waves + +Two tickets that both edit the same file are a real merge-conflict risk when run in +parallel isolated worktrees. Group tickets into waves: +- Tickets with no file overlap with each other → same wave, dispatched in parallel. +- A ticket that heavily edits a file another ticket already changed → its own later + wave, dispatched only after the earlier one has merged (so its worktree branches from + a base that already has the earlier change). +- Light, disjoint-region overlap in the same file (e.g. two tickets touching different + methods of the same screen) is an acceptable same-wave risk — it produced a small, + manually-resolvable conflict in practice, not a blocker. + +Write this wave plan into the ticket README before dispatching. + +## Phase 5 — Dispatch, implement, merge (per wave) + +For each ticket in the current wave, launch an `Agent` call with `isolation: "worktree"` +and the **full ticket file contents inlined in the prompt** (not just a path — a fresh +agent should not have to go re-discover context this phase already gathered). Write the +prompt itself in ASD-STE100 too — the dispatch prompt is the second place, after the +ticket file, where ambiguity costs the most. The prompt must tell the subagent to: + +1. Read the inlined ticket, implement exactly what its Design/Implementation sections + specify — it's already been reviewed, not something to redesign from scratch. +2. Run `flutter analyze` (must stay clean) and `flutter test` (must stay green, count + must not regress below the baseline from Phase 0/the previous wave). +3. Append an `## Outcome` section to its own ticket file: what shipped, any deviation + from the design and why, the final test count. If an assumption in the ticket turns + out wrong once real code is touched, say so honestly here rather than silently + patching over it. +4. Flip its own ticket's status line to Done — but **not** the shared README's table; + that gets updated once after merging, by you, not by parallel subagents racing each + other on the same file. +5. Commit locally. Follow this repo's normal commit convention unless the user asked + for something different for this run (e.g. no co-author trailer) — confirm which + applies during Phase 2 if it's ambiguous, don't assume silently either way. +6. State explicitly that Android-emulator verification is **best-effort only** at this + stage (Phase 6 does the real device pass) — a subagent that gets stuck fighting + emulator flakiness instead of finishing its actual ticket has failed the ticket. + +Wait for each subagent in the wave to report back, then merge its branch: + +```bash +git merge --no-ff worktree-agent-<id> -m "Merge <ticket-id>: <title>" +``` + +A conflict in the ticket `.md` file itself (both sides changed the status/Outcome +header) is expected and mechanical — take the incoming side +(`git checkout --theirs docs/<slug>/<ticket>.md && git add ...`), since it carries the +Outcome section. A conflict in real code needs an actual read of both diffs — do not +blindly pick one side; understand what each branch was trying to do and merge the +*intent*, not just the text (see the worked example: two tickets each added a different, +correct enhancement to the same function, and the right merge combined both rather than +picking one). + +After every merge: `flutter analyze` clean, `flutter test` green, count check. Only then +move to the next ticket in the wave or the next wave. Update the ticket README's status +table to Done for everything just merged, as its own small commit. + +## Phase 6 — Emulator verification pass (all waves merged) + +This is not optional and not the same thing as the best-effort check individual +subagents did — this is the real bar. Boot (or reuse, if already running and +responsive) the project's emulator, launch the freshly-merged app, and walk every +ticket's golden path for real, taking screenshots as evidence. + +**Known failure modes from prior runs, and how to handle them without losing an hour:** +- **Coordinate scaling.** Screenshots are shown to you scaled down (e.g. displayed at + 900×2000 for a native 1080×2400 device) — the image caption states the exact ratio. + Multiply *both* x and y by that ratio before issuing `adb shell input tap`. Getting + this wrong looks like "the app isn't responding to taps" and wastes real time — it + isn't a bug, it's a unit error. When in doubt, pixel-sample the raw screenshot file + directly (PIL/`Pillow` in Python) to find a real on-screen coordinate rather than + eyeballing the scaled-down preview. +- **ANR / "isn't responding" loops.** Check host memory pressure first + (`vm_stat`, `sysctl vm.swapusage`) before assuming an app bug — a memory-starved host + produces exactly this symptom. Escalate in order: dismiss-and-wait once → force-stop + and relaunch the app (`adb shell am force-stop <pkg>` then `am start`) → full emulator + restart (`adb emu kill` + relaunch) → if the host itself is out of memory, say so and + ask the user to free some rather than fighting it silently for another hour. + `com.rippr.port` / `com.rippr.MainActivity` are this app's package/activity for direct + `am start` (faster than a full `flutter run` rebuild when the APK is already + installed and you just need a clean process). +- **Airplane mode / no network** on a fresh emulator boot makes map tiles fail to load + (renders as a checkerboard, not the app's real tiles) — check + `adb shell settings get global airplane_mode_on` and + `adb shell cmd connectivity airplane-mode disable` before concluding tiles are broken. +- **Stale app state** (a leftover active recording, an old database) from a previous + manual test session can make a screen look wrong for reasons unrelated to the current + tickets — `adb shell pm clear <pkg>` before a verification pass if there's any doubt + the state is clean. +- **Location-dependent features** (anything reading GPS) need a granted permission and + a mock fix on the emulator: `adb shell pm grant <pkg> android.permission.ACCESS_FINE_LOCATION` + (and `ACCESS_COARSE_LOCATION`), then `adb emu geo fix <lon> <lat>` — note longitude + first. A feature that reads location may take several seconds after granting/fixing + before the first real fix arrives; don't conclude it's broken from a screenshot taken + immediately after. + +**When the emulator pass finds a real bug** (as opposed to an environment artifact +above): fix it directly, add a regression test that fails without the fix and passes +with it, re-run `flutter analyze`/`flutter test`, commit with a message that says what +the emulator pass found and why the widget tests missed it, then re-verify on-device +that the fix actually shows correctly. Do not mark a ticket's on-device verification +done until you've seen the fixed behavior with your own eyes (a screenshot), not just +inferred it from the test passing. + +If the emulator is genuinely too unstable to get a real signal (host resource +exhaustion that persists after the escalation steps above), say so plainly rather than +claiming verification that didn't happen — this exact honesty call was made once +already in this repo's history and is the correct call to make again if needed. + +## Phase 7 — Final commit: fresh APK + +Once every ticket is merged, green, and emulator-verified: + +```bash +flutter build apk --release +git add -f build/app/outputs/flutter-apk/app-release.apk +git commit -m "<summary of what this batch shipped>" +``` + +Force-add is required — `build/` is gitignored, and this repo's own established +precedent (see `git log --oneline -- '*.apk'`) is to force-add the release APK directly +into the repo rather than relying on CI artifacts or a release page, since there is no +CI/release pipeline set up. Match that precedent unless the user has since set one up. + +## Phase 8 — Report back + +Summarize: what shipped (ticket list with one line each), test count before → after, +what the emulator pass confirmed (and any bugs it caught that the unit-test bar alone +would have missed), and the final commit. State plainly whether anything was pushed +(this skill never pushes on its own — say so, and say what's ready for the user to push +if they want to). + +## What this skill is not for + +A single small, obviously-scoped change (one file, one clear fix) doesn't need five +phases of ceremony — just make the change, test it, done. This skill is for a +requirements document with enough surface area that breaking it into independently +verifiable pieces is itself valuable — the kind of input that would otherwise become one +enormous, hard-to-review commit. diff --git a/.gitignore b/.gitignore index b637bf0..fbba1b1 100644 --- a/.gitignore +++ b/.gitignore @@ -50,4 +50,4 @@ app.*.map.json # Build artifacts (large; regenerable) build/ .dart_tool/ -.claude/ +.claude/worktrees/